Thursday, September 22, 2011

Postdoc in phyloinformatics available

A postdoc position in phyloinformatics is available at the Florida Museum of Natural History, University of Florida, to work on the integration of database developments and analytical workflows.  Experience with phylogenetic and workflow software (RaXML, MrBayes, Galaxy, Kepler etc.) and a background in data management is highly desirableThe successful candidate must have a PhD in biology, computer science, computational biology or related fields. Additionally, Java will most likely be your everyday cup of coffee.  The position is available for at least 2 years. Please contact me for additional details.

Wednesday, September 7, 2011

Lab openings

I have openings in my lab to work on the evolution, systematics and biogeography of preferably Campanulaceae and Melastomataceae (but I am open to other groups too).  If you are interested, contact me directly and apply to UF by the 31st of December, 2011.

Saturday, August 27, 2011

Call For Participation: Steps towards a Minimum Information About a Phylogenetic Analysis (MIAPA) Standard

Synopsis
Many phylogenetic analysis results are published in ways that present serious barriers to their reuse in numerous research applications that would stand to benefit from them. While some of these barriers are well understood, such as issues with adherence to standard exchange formats, those centering on the associated metadata necessary for researchers to evaluate or reuse a published phylogeny have only recently begun to be articulated. One of the critical next steps towards formalizing these metadata requirements as a minimum reporting standard is to convene meetings of key stakeholder communities with the goal to identify information attributes  necessary and desirable for facilitating reuse, and to build consensus on their priority. To this end, we are holding a workshop at the 2011 Biodiversity Information Standards (TDWG) Conference to determine how a future reporting standard for phylogenetic analyses can best serve biodiversity science and related research applications.  We invite all interested colleagues to participate.

Background

The workshop of the Biodiversity Information Standards (TDWG) Phylogenetics Standards Interest Group held at the 2010 TDWG conference included a project focused on how to publish re-usable trees that can be linked into an emerging global web of data.  Through follow-up work, this led to the following tangible results:
  1. An online draft report of the 2010 TDWG workshop [1], and a corresponding manuscript on best practices for publishing phylogenetic trees (Stoltzfus et al. in preparation);
  2. An 2011 iEvoBio presentation on “Publishing re-usable phylogenetic trees, in theory and in practice” [2];
  3. lighting talk presentation and Birds-of-a-Feather gathering at 2011 iEvoBio, and
  4. A survey group that explored barriers to re-use and developed plans for a survey
These activities have considerably clarified our understanding of the theory and practice of publishing re-usable phylogenetic trees: how many phylogenies are published each year, the (low) frequency of archiving, what archives and tools are available, what policies are in force, etc.  We have identified a number of barriers to re-use involving such aspects as technology, standards, culture, and access.  
Many of these barriers can be interpreted as a consequence of the lack of a community-agreed standard for what constitutes a well documented phylogenetic record.  In the absence of such a standard, trees are often archived as image files rather than in appropriate data exchange formats, and lack important accompanying information (metadata), such as externally meaningful identifiers, that would be needed to make them useful to others. The idea of a Minimum Information About a Phylogenetic Analysis (MIAPA) standard has been suggested [3], but so far there has not been a deliberate process to develop and disseminate a community standard.  Meanwhile, a number of systematics and evolution journals have begun to require archiving of the data underlying published research findings [4].  The emerging cultural shift in data archiving and sharing promoted by this policy change offers a unique window of opportunity to move ahead with the development and actual specification of a MIAPA standard.
Similar to other minimum reporting standards [5], the primary focus of a future MIAPA standard would be on defining a “checklist” of metadata information attributes that, at a minimum, needs to accompany an archived phylogenetic analysis, and to which standards values for these attributes would need to adhere. The key step in developing community consensus on these elements of the standard is to convene a series of meetings that collectively involve participants from all major groups of stakeholders who would be affected by such a standard, such as users, producers, publishers, or archivists of phylogenetic analyses.  To aid this process, the Phylogenetics Standards Interest Group is holding a workshop at the 2011 TDWG conference, with the goal to obtain consensus requirements and priorities for a MIAPA checklist for the purposes of biodiversity science, taxonomy, museum collections, and related research applications.

Goals and deliverables

The main goal of the workshop is to develop a shared understanding of the role that a MIAPA standard could play in facilitating re-use of phylogenetic analyses for the biodiversity science and related communities, and what the standard would need to specify in order to  best fill that role. Possible deliverables include
  1. A draft set of information attributes that should or could be included in a provisional MIAPA checklist, with a level of consensus for each of them.
  2. A database with use-cases based on exemplifying publications, that report phylogenies to elucidate a broad spectrum of questions relating to biodiversity science.
  3. A refined MIAPA survey to be informed by biodiversity science cases for reuse.
  4. A plan for further community engagement and consensus-building among biodiversity science stakeholders.

Workshop format

The workshop will start with a few presentations focused on (i) introducing MIAPA and its potential in facilitating reuse (J. Leebens-Mack); (ii) summarizing recent developments and current status of MIAPA-related efforts (A. Stoltzfus); and (iii) past experiences and resulting best practice recommendations on developing a minimum reporting checklist standard (D. Field). The rest of the workshop will be hands-on.  Participants in the workshop will break out into groups to address separate issues according to the anticipated deliverables and best practice recommendations.
The workshop will be 1.5 days in duration, and be held during the 2011 Biodiversity Information Standards (TDWG) conference, to take place Oct 17 to 21, 2011 in New Orleans, USA. (http://www.tdwg.org/conference2011/).  The workshop will start in the afternoon of Monday, Oct 17, and end on Tuesday. Oct 18.

How to participate

Participation in the workshop is open to everyone interested. However, space is limited, and we therefore ask that, if you are interested in attending, to please communicate your interest through the MIAPA discussion group [6]. This will also allow us to include you in pre-workshop planning. Since the workshop is part of the TDWG conference, participants will need to register either for the full conference, or for the days of the workshop.  
The organizers will provide an electronic venue for participants to share ideas and develop plans in advance of the workshop.  After the initial presentations, participants will self-organize into task groups.  
Organizers
  1. Nico Celinese, University of Florida
  2. Hilmar Lapp, NESCent  
  3. Jim Leebens-Mack, University of Georgia
  4. Enrico Pontelli, New Mexico State University
  5. Arlin Stoltzfus, NIST & University of Maryland

References

[1] Whitacre et al. (2010). Current Best Practices for Publishing Trees Electronically. http://wiki.tdwg.org/twiki/bin/view/Phylogenetics/LinkingTrees2010
[2] O’Meara et al. (2011). Publishing re-usable phylogenetic trees, in theory and practice. Available from Nature Precedings<http://dx.doi.org/10.1038/npre.2011.6048.1>
[3] Leebens-Mack, J., T. Vision, et al. (2006). "Taking the first steps towards a standard for reporting on phylogenies: Minimum Information About a Phylogenetic Analysis (MIAPA)." Omics 10(2): 231-7.
[4] Whitlock, M., M. McPeek, M. Rausher, L. Rieseberg, and A. Moore (2010). Data Archiving (Editorial). The American Naturalist 175(2): 145.
[5] Taylor, C.F., D. Field, S. Sansone, J. Aerts, R. Apweiler, M. Ashburner, C.A. Ball, et al. (2008). Promoting coherent minimum reporting guidelines for biological and biomedical investigations: the MIBBI project. Nature Biotechnology 26(8): 889-96. doi:10.1038/nbt.1411
[6] MIAPA discussion group: http://groups.google.com/group/miapa-discuss

Thursday, July 21, 2011

Why am I going to TDWG 2011?


Well, I promise you, the location has little to do with it (well, maybe more than just a little). I already envision myself sitting in some cool bar in the Frech quarter, eating soul food (drinking a little), listening to great music and talking about, let's see, biological collections digitization, data acquisition, data integration, phylogenetics standards, interoperability, cool new tools (like BiSciCol...because I am not biased :-), etc. etc.

Jokes aside, why do I really want to go to TDWG this year?!
There is so much going on right now, such a renewed interest in collections and digital data. First BiSciCol and then VertNet were funded, iDigBio is in place, 3 awesome Thematic Collections Networks have also been funded, covering 90+ Institutions in 45 US States, in addition to a bunch of other collections being supported by NSF through their regular programs. Can it get much better in these economic times? Some exciting new blogs have bee popping up lately, proposing cool ideas, different approaches, encouraging us to think outside the box.  Yes, we've been talking about these topics for a long time but I get a clear sense that now we can DO things, and can go well beyond talking about them.  We can actually experiment now and scale up our ideas to see if new approaches can be successfully implemented. We finally have the means! We have the attention of our funding agencies and a few seeds have been planted already (can't stop botanizing!).

I am excited because with this renewed feeling of being able to actually change things, make progress, provide a clear input, we can come together as an inclusive community.  This is perhaps a first solid opportunity for real cross-fertilization among different groups, like TDWG and SPNHC, BiSciCol with VertNet and other domain specific networks, and what about getting more biologists involved in the geek world?  Not that we haven't been supporting each other before, but I can see this time is different.  This time is not just about the momentum, the charge we all feel when we get together at a meeting and plan ahead. This time we can all be excited about going back home and getting down to work! I am really hopeful that the recent events and investment provide the glue that we all have been needing for a long time. We are in the same boat and everyone's input is no little contribution.

We have an opportunity to work as a community, think globally and act locally (who said that?! Feels nice right now!). It is not anyone's mission to succeed, it is OUR collective mission.  So, I am excited to gather around our common problems and bottlenecks, and being able to concretely share the load by developing parallel approaches and putting them to work, together. The goal ultimately is to create an environment where we can all do better science, cool science, where posing new challenging questions will be fun because we know we will have the means to answer them.  It's going to be a great playground! And that's why I am going!

Friday, June 24, 2011

My post-iEvoBio Meeting emotional outburst

I have just come back from the iEvoBio Meeting in Norman, Oklahoma. This is the first time ever I didn't quite mind to be stuck in a place in the middle of nowhere (seriously!) because my fellow prisoners were actually pretty entertaining.  This is what I love about iEvoBio. It does NOT bore you and it keeps you awake despite the heavy drinking session of the previous night; and for someone like me, with ADD, the constant flow of information, often delivered in 5 minutes slots, is just perfect (it's like watching great superbowl-style commercials!). The meeting format offered a little bit for every taste, from more focused presentations to fairly short discussion sessions. What do we get from short discussion sessions? Well, how about a set of quickly vomited ideas, needs, wish-lists, from a variety of people with often different backgrounds, seeds to bring home and plant to see whether anything germinates, either in your lab or someone else's. This is what's fun in science! I just love to be there and watch what will happen next, e.g. next year at iEvoBio.  We really need iEvoBio to remind us every year of the status quo in evolutionary informatics, what's cooking, what are the missing ingredients, the limitation and potential of the tools we build, and the beauty is we don't really need a week to catch-up.  Two intense days (and nights :-) are actually great!  Last year in Portland, OR, the meeting was fantastic, but this year I enjoyed it no less! For a great list of projects and new tools presented see Recology blogpost. Just a few of my favorites include Map of Life, TreeBASE with an R interface, Phenoscape, Ontogrator, and BirdVis. Of course, we presented our BiSciCol prototype and our slides are up in slideshare now. We didn't win the challenge though, despite me trying to cheat the system by voting from different browsers (I couldn't help it, it was more of an experiment, really ;-) but BiSciCol is still in its infancy stage and we will present a more mature product at the next TDWG meeting in New Orleans, so stay tuned!
Can't wait for next iEvoBio meeting in Ottawa! In the meanwhile I am charged-up and super ready to go back to work!

Wednesday, April 13, 2011

Talking about names....

A new Angiosperm phylogeny paper is finally out!  It clarifies relationships at some of the deeper messier nodes and reinforces our previous knowledge of some others.  The point I want to make here is that the authors use phylogenetic nomenclature to name major clades.  They state:

"For higher clades, we consistently use PhyloCode names (see Cantino et al., 2007 ) whenever these are available; these names are always in italics (e.g., Pentapetalae, Mesangiopsermae, Rosidae, Fabidae, Malvidae). Note that Rosidae (sensu Cantino et al., 2007 ) does include Vitaceae. Our use of family and ordinal names follows APG III (2009) as a formal point of reference; for Caryophyllales, we follow Cantino et al. (2007; hence, the use of italics), which matches the APG III circumscription. For additional recent discussion on families and their status, see the Angiosperm Phylogeny Website (Stevens, 2001 onward). We recognize that some broader family circumscriptions favored in APG III are controversial and can obscure underlying diversity (e.g., Passifloraceae s.l.), which would be evident with narrower circumscriptions."

I hear left and right how the PhyloCode is going to bring a mess in the field of Biology by changing all names and reshaping the Classification (impossible, as it is just a nomenclatural tool). Well, here it is, another shocking paper that imposes so many new names to the community. The way I see it, we now have some pretty well refined concepts attached to names, the same names that have been previously used idiosyncratically. Names that won't change even if the clade content does change. Names that will always refer to the same ancestors. Names that we could actually query and be happy with what we retrieve. That would be indeed a breath of fresh air.