Showing posts with label Workshop. Show all posts
Showing posts with label Workshop. Show all posts

Wednesday, August 22, 2018

Distinguishability in Phylogenetic Networks, report


We have now completed the workshop, as you can tell from the previous post with some photos. Here is a brief report on what seem to me to be some of the more useful points covered.


We had 10 formal presentations, but we also focused on group discussions for several hours each day. It may be the latter that were the most productive. However, I will briefly summarize the talks first.

I spent my time time in the opening talk emphasizing the different viewpoints of network computations, which focus on the patterns that can be detected in the data, and the network users, who are generally more interested in the processes that create those patterns (or are, indeed, absence from the patterns but present in the phylogenetic history, anyway). This highlights the two essential point of the workshop title, that both the patterns and the processes are much harder to untangle for networks than for trees.

Céline Scornavacca then bravely tried to tackle the combined problem, anyway, by trying to produce networks from analyzing the patterns in terms of their processes. The issues immediately become obvious, but she seems to be determined to proceed, regardless. Later in the week, Luay Nakhleh reduced the issue simply to vertical processes (including incomplete lineage sorting but not gene duplication-loss) versus horizontal processes. This creates a tractable problem for parsimony and likelihood, but the current challenge remains the limited number of taxa.

Vincent Moulton, Cécile Ané and Charles Semple dodged the issue by focusing on computations. Charles took on the challenge of trying to create a network version of Neighbor-Joining, which would address the issues of computational speed and taxon sampling, and Vince tackled super-networks, and the conditions required for building networks from a collection of smaller (ie. incomplete) trees. Both topics remain open questions. Cécile, on the other hand, discussed network models for trait evolution, which is important for the use of phylogenetic comparative methods when using networks.

On the user side, the presentations focused on examples, and the issues encountered when dealing with them. James Whitfield and Axel Janke talking about biology (mostly phylogenomics), while Johann-Mattis List talked about linguistics, and Tiago Tresoldi talked about stemmatology. In some ways, historical linguistics seems to be the odd one out, since many of the issues dealt with are somewhat removed from those in the other fields. However, in biology there are actually two options for producing networks — directly from the data or via "gene trees" (trees derived from non-recombining blocks of sequences). For the humanities, much of the current discussion is about the nature of the data, and how to code it for quantitative analysis.

This brings us to the discussions. While some time was spent on trying to establish whether biologists think that there is a difference between lateral gene transfer and horizontal gene transfer, or between incomplete lineage sorting, ancestral polymorphism and deep coalescence, some productive interchanges also occurred. Here is a coverage of four of the most important ones.

There was general agreement that there are several barriers to widespread adoption of network analyses in phylogenetics. This includes the development of suitable methods (in the face on indistinguishability), but also includes an understanding of what methods are currently available, what data are required to apply those methods, what taxon sampling is required to benefit from the methods, and how to use the programs that implement those methods.

One popular suggestion was therefore to produce some sort of "cookbook", to address the complexity of producing networks, given that there are many methods and programs. From the users' point of view this would illustrate what network analyses can do, in terms of finding reticulation patterns in the data; and from the computational point of view it would outline what needs to be done to get the programs to work. The consensus idea was to choose two suitable datasets (yet to be determined), and then have each program author provide analyses of them (including any scripts that are needed).

Following on from this latter point, it was agreed that the programs need easy user interfaces, if they are to become more widely used. Here, the word "widely" includes casual users from outside of phylogenetics, who use phylogenies as only one of many tools in their work. So, users will include those who need nothing more than a "point and click" control panel (which may be >90% of potential users) to those who would benefit from scripting control of the analyses. The interface needs both a front end, to specify the particular analysis, and a back end, to allow exploration of the output.

Another long-discussed issue was how to popularize networks, which is clearly a major topic. A phylogenetic tree is nothing more than one of the possible networks for any given dataset, and yet the focus is often on trees rather than networks.

To this end, it was noted that the current Wikipedia entry is inadequate, especially compared to the corresponding entry for phylogenetic trees. Not only is this entry out of date, it is in a number of ways misleading. In particular, there needs to be a discussion of the fact that, if a network is a "tree with reticulations", then ignoring the reticulations can result in the wrong tree, and the branch lengths may be severely under-estimated. There are challenges to getting Wikipedia entries changed, especially the wholesale re-writing of an entry, but this will be necessary.

Finally, it was noted that Philippe Gambette's Who is Who in Phylogenetic Networks website is extremely useful but is still poorly known, even within the phylogenetic networks community. We had a long discussion about how to enhance this site, to make it a more general-purpose repository of information about phylogenetic networks. This included a more inclusive database, more comprehensive tagging of keywords, enhanced descriptions of those keywords, and ways to keep the database up to date.


Steven Kelk has the notes from the final session, which was a review of what we achieved during the workshop, and which contains the To Do list. Both he and Philippe have the notes about modifications for the Who is Who in Phylogenetic Networks website, which is likely to be the first outcome-project tackled.

Thankyou to everybody who participated in the workshop. It seemed to be very productive, with a number of concrete outcomes that will be interesting to review at the next workshop.

Friday, August 17, 2018

Distinguishability in Phylogenetic Networks, photos


Evidence that we were in the Netherlands.



Evidence that we did some work.



Left to right: Steven Kelk, David Morrison, Mike Steel, Philippe Gambette (obscured), Tiago Tresoldi, Claudia Solis-Lemus, Fabio Pardi, Simone Linz, Mark Jones.


Left to right: David Morrison, Cecile Ané, Philippe Gambette (obscured), Katharina Huber, Leen Stougie, Remie Janssen, Yukihiro Murakami, Mattis List, Gereon Kaiping and Charles Semple.


Left to right: David Morrison (obscured), Axel Janke, Steven Kelk, Charles Semple, Claudia Solis-Lemus, Mark Jones (obscured), Fabio Pardi, Leo van Iersel, Simone Linz and Vincent Moulton.


Céline Scornavacca lectures Cecile Ané.


Axel Janke and Leo van Iersel contemplate methods for infering hybridization.


Philippe Gambette and Guido Grimm.


Mozes Blom and Jim Whitfield.


Mike Steel and Luay Nakhleh.


Luay delivers his Final Message, to Mozes Blom, Cecile Ané, Katharina Huber and Charles Semple.


Monday, November 9, 2015

Capturing phylogenetic algorithms for linguistics


A little over a week ago I was at a workshop "Capturing phylogenetic algorithms for linguistics" at the Lorentz Centre in Leiden (NL). This is, as some of you will recall, the venue that hosted two earlier workshops on phylogenetic networks in 2012 and 2014.

I had been invited to participate and to give a talk and I chose to discuss the possible relevance of phylogenetic networks (as opposed to phylogenetic trees) for linguistics. (My talk is here). This turned out to be a good choice because, although phylogenetic trees are now a firmly established part of contemporary linguistics, networks are much less prominent. Data-display networks (which visualize incongruence in a data-set, but do not model the genealogical processs that gave rise to it) have found their way into some linguistic publications, and a number of the presentations earlier in the week showed various flavours of split networks. However, the idea of constructing "evolutionary" phylogenetic networks - e.g. modeling linguistic analogues of horizontal gene transfer - has not yet gained much traction in the field. In many senses this is not surprising, since tools for constructing evolutionary phylogenetic networks in biology are not yet widely used, either. As in biology, much of the reticence concerning these tools stems from uncertainty about whether models for reticulate evolution are sufficiently mature to be used 'out of the box'.

As far as this blog is concerned the relevant word in linguistics is 'borrowing'. My lay-man interpretation of this is that it denotes the process whereby words or terms are transferred horizontally from one language to another. (Mattis, feel free to correct me...) There were many discussions of how this proces can confound the inference of concept and language trees, but other than it being a problem there was not a lot a said about how to deal with it methodologically (or model it). One of the issues, I think, is that linguists are nervous about the interface between micro and macro levels of evolution and at what scale of (language) evolution horizontal events could and should be modelled. To cite a biological analogue: if you study populations at the most microscopic level evolution is usually reticulate (because of e.g. meiotic recombination) but at the macro level large parts of mammalian evolution are uncontroversially tree-like. In this sense whether reticulate events are modeled depends on the event itself and the scale of the phylogenetic model concerned.

Are there analogues of population-genetic phenomena in linguistics, and are they foundations for phenomena observed at the macro level? Is there a risk of over-stating the parallels with biology? One participant told me that, while she felt that there was definitely scope for incrorporating analogies of species and gene trees within linguistics - and many of the participants immediately recognized these concepts - comparisons quickly break down at more micro levels of evolution.

I'm not the right person to comment on this of course, or to answer these questions, but in any case it's clear that linguistics has plenty of scope for continuing the horizontal/vertical discussions that have already been with us for many years in biology...

Last, but not least: it was a very enjoyable workshop and I'm grateful to the organizers for inviting me!

Friday, July 31, 2015

Singapore, Day 5

Today was a mixed bag ot talks.

Louxin Zhang started with a couple of proofs about what he called "stable" networks; and Stefan Grünewald developed his thoughts on quartet algoritms for splits graphs. At the other extreme, Nadine Ziemert talked entirely biology, introducing the audience to the problem of trying to study the evolution of secondary metabolites. In between, Eric Tannier tried to use horizontal gene transfer to date the nodes of networks, assuming that HGT requires a temporally consistent network. Francois-Joseph Lapointe produced the only really statistical talk of the week, trying to produce p-values for patterns on sequence similarity networks.

Daniel Huson popped in for the last day, and presented us with some ideas for the future development of both SplitsTree (unrooted networks) and Dendroscope (rooted networks). Apparently, the need is for SplitsTree to handle larger sets of trees, while for Dendroscope it is to produce networks from pairs of input trees. He also noted that there are still more networks being produced using median joining rather than neighbor-net, due to the amount of work being done on human mitochondrial sequences.

An interest was expressed in continuing the series of meetings on phylogenetic networks (Leiden 2012, Leiden 2014) — I first met most of the people working on networks in phylogenetics in Uppsala in 2004 (Phylogenetic Combinatorics and Applications).

Today we also celebrated Dan Gusfield's 2^6 birthday, with a strawberry cream cake.

So, all in all, a very successful meeting.

After the sessions finished, I went down to the Gardens By The Bay to look at the Supertree Grove. As you can see, a "super" tree is by any definition actually a network.


Thursday, July 30, 2015

Singapore, Day 4


There was more heavy maths today.

Charles Semple started by counting trees within specified types of network. In the process, he provided the first mathematical proof of the week (he actually provided two). He also raised the issue of what, exactly, is a phylogenetic network — we have had many mathematical restrictions placed on networks this week, and it is not always clear how any of them might relate to biological concepts.

Leo van Iersel tried constructing super-networks from incomplete sub-networks, sticking to algorithms rather than proofs. Yufeng Wu and Zhi-Zhong Chen later tried the same strategy for their networks, as did Lusheng Wang for pedigree comparison (he was the only person other than myself to even mention pedigrees).

Mike Steel considered under what circumstances a network can be viewed as a "tree with reticulations" rather than a non-tree network (ie. not every vertex is part of the same underlying tree); this led him to the interesting observation that whether a dataset can be represented by a tree can depend on the taxon sampling. He also looked at when a set of non-tree distances can appear to be tree-like, which is the sort of question that only a mathematician would ask.

Most of the audience interjections this week have come from Sagi Snir, and the rest of the speakers got to return the favor this afternoon, when he spoke about trying to reconstruct trees subject to large amounts of horizontal gene transfer. In the process, he also tried to "sketch" a mathematical proof, which turned into a full-sized painting, before moving on to his algorithm.


Wednesday, July 29, 2015

Singapore, Day 3


It was hard going this morning for the biologists, as there were three main computational talks. First, Vince Moulton further developed some of his ideas about split networks, including median networks, quasi-median networks and neighbor-nets, and what sorts of trees they might contain. Then, Céline Scornavacca expanded on her ideas for calculating the "hybrid number". Finally, Jens Lagergren outlined his work on fitting gene trees to known species trees and networks; this has come a long way in recent years.

We had the afternoon off, although many people took the opportunity to pretend that they were still in their offices at home. Myself, I sat by the pool waiting for the temperature to cool (this was the hottest day so far this week), and then went to the Singapore Botanical Gardens, where I circumnavigated the Evolution Garden, the National Orchid Garden, and the Rainforest Walk. I then briefly perused Orchard Street (one of the most ridiculous shopping meccas you will ever see) and the Raffles Hotel (an even more ridiculous hang-over from British imperialism), before returning to the pool side. It's a tough life.


Tuesday, July 28, 2015

Singapore, Day 2

The computational people were very patient today, as the three major talks focussed on biology, with only the shorter talks being computational.

In particular, Eric Bapteste and James McInerney were determined to tackle the true complexity of phylogenetics, rather than trying to see genealogical history as being a tree with reticulations. They have recently been championing sequence similarity networks as tools for exploring phylogenetic history, and Eric discussed them in relation to prokaryote evolution while James looked at gene families. Strictly speaking, SSNs are not phylogenetic networks, because they do not involve the inference of unobserved nodes connected to observed (labelled) nodes by inferred edges, bit instead connect observed (labelled) nodes via observed edges. This does not mean that they have no role to play in phylogenetics, as the speakers made amply clear.

My own talk had little to do directly with empirical networks, but instead tried to look at an overview of the field, presenting some of my own ideas about where networks are heading, and what role they might play as phylogenetic tools. Not everyone was convinced.

Also of personal interest to me, Philippe Gambette unveiled the new, much more ambitious, version of the Who is Who in Phylogenetic Networks database. I claim no other role in this than encouraging Philippe to be as ambitious as possible. I think that people will be genuinely impressed by what can now be done to explore the people, literature and software associated with phylogenetic networks.

PDF copies of the talks have now started to appear on the workshop web page, which will give you a bit more idea of what our speakers have tried to say.

We also started the discussion about how to engender more effective development of computational tools for phylogenetic networks. Topics covered included the need for more gold-standard datasets that can be used to test new methods — to date, the ones available on this blog have been compiled by me alone, but in order to expand this other people will need to contribute. Also, improved communication and collaboration between biologists and computationalists would be very helpful, and several suggestions were canvassed, but no real way forward was found. One interesting point was made that many of the practical applications of networks were not likely to attract the professional interest of most computationalists — indeed, to date, most phylogenetics programs have been written by biologists rather than computational people.

In the evening, I had a very enjoyable dinner at the Singapore Seafood Republic, in the company of Louxin Zhang (our organizer) and some very nice Chinese visitors, all of whom politely failed to comment on my inability to use chopsticks. Pictures of dinner were taken, and may appear on Facebook, if I am not careful. We finished just in time to watch the Sentosa Crane Dance, which you should all check out.

Monday, July 27, 2015

Singapore, Day 1


I'm not sure how these reports are going to go, as I did not bring a laptop with me. Also, Blogger is not happy with me logging in from another country. However, I have managed to get a decent sized keyboard on the screen of my iPad Mini, so I can at least type somewhat normally. I will not, however, write about every talk (and my apologies to those speakers who do not get mentioned).

Singapore is as expected — hot and humid; except when one is indoors, and even 24 °C surprisingly feels cold. I have washed most of my shirts once already, to remove the perspiration.

Most people seem to have arrived; indeed, many have already been here for a few days. Myself, I spent Sunday afternoon touring Chinatown, and Little India. The food market at the latter location was unbelievably hot, although the locals did not seem to realize this.

We have now dealt with the first day of talks. No mercy was shown to the uninitiated, and we started with the heavy network stuff right from the start.

This took the form of Dan Gusfeld explaining to us in no uncertain terms that Integer Linear Programming can be used to solve many computational problems that are too hard for Dynamic Programming, using Ancestral Recombination Graphs as his example. When asked about possible connections to actual biology, he patiently explained that this was another matter entirely. Kathi Huber later said the same thing when asked about the loss of biological information resulting from unrooting a rooted network. At the time, she was trying to "bridge the gap" between rooted and unrooted networks, and unrooting them is surprisingly effective way to achieve this.

Luay Nakhleh's talk was my favorite of the day. He is one of the few people in this business who can successfully talk computations to a mathematician and biology to a biologist — most of the rest of us fail at one or the other (or both). Sadly, he pointed out that under the coalescent model any gene tree fits inside any species tree (or network), simply by having the gene coalescences occur after the species root is reached. He also noted that we can't distinguish among reticulation processes on a network, which took away one third of my talk!

We finnished with Jesper Jansson decomposing networks into triangles, which is a neat change from the usual decomposition into triplets, clusters or trees. Along the way, he concluded that we need to keep using a lot of different measures for network to network distances, because none of the current ones are good under all conditions. That is another major difference between trees and networks.

Wednesday, July 22, 2015

Phylogenetic Network Workshop, Singapore


Next week there will be a gathering in Singapore, for a Phylogenetic Network Workshop. This is being hosted by the Institute for Mathematical Sciences, at the National University of Singapore.


The workshop has been organized under the guidance of Louxin Zhang. The program and abstracts can be found here. It runs for the whole week, 27 – 31 July 2015.

The workshop is actually the final part of a much larger, 2-month programme, called Networks in Biological Sciences (1 June – 31 July 2015). This programme is focused on mathematics for network models in biology, including complex networks and systems biology. Network modeling is extremely challenging, and so it offers outstanding opportunities for mathematicians and statisticians. The phylogenetics workshop will focus on the mathematics needed to develop fast and robust computer programs for inferring an evolutionary network models from biological sequence data.

The participants are principally from the computational sciences, of course, including many who have attended the previous network workshops in Leiden, in the Netherlands, in October 2012 and July 2014. There are, however, a few biologists to round out the field, including myself.

Singapore is hot and humid for most of the year, and July is no exception. So, I am expecting the unacclimatized participants to spend most of their time indoors, avoiding the daily thunderstorms.

I am hoping to add some blog posts based on what happens at the workshop, as it proceeds.

Wednesday, July 16, 2014

Touching the Data, photos

We all worked hard during the workshop. Here is our fearless leader, in deep thought:


While some of the younger participants enjoyed drawing on the walls:


Professor Whitfield has come up with a great new model of evolution: phylogenetic windmills:



There was not only work, but also time to relax and enjoy the beautiful Dutch summer weather:


And not to forget the delicious Dutch food:


But really, most of the time we were busy touching the data, which you can find on this website:


For more photos, see the Touching the Data website.

Friday, July 11, 2014

Touching the Data, report 2


We have now completed the workshop.

Since the first report, we have had three more talks. First, Mukul Bansal outlined the relationship between phylogenetic networks and reconciliation analysis, and the way in which the latter can be used to construct the former. Starting from an estimated species tree, the tree for each locus is optimized for fit to the species tree, which helps locate any areas of extensive gene flow (ie. reticulation). This can be done using a large number of loci and an even larger number of taxa.

Celine Scornavacca provided details of some of the fundamental limitations of network analysis.The most important of these is unidentifiability of network topologies -- there are classes of network topologies that cannot be distinguished based on the information that is currently used, so that we cannot guarantee that a unique optimal network will be found during an analysis. Branch lengths may help with this situation, but cannot guarantee to resolve it.

Jim Whitfield covered the advantages and potential problems of using genomic-scale data for phylogenetic analysis. The basic problem is the increased scope for error in moving to the genome data (genome assembly problems, gene homology issues, alignment difficulties), although the potential advantages are extensive.

Most importantly, we spent two days "touching" some data. The participants broke into smaller groups of continuously varying size, each of which focussed on a particular dataset (as supplied by some of the participants). These data were evaluated in many different ways, to assess the characteristics of the data as well as to evaluate the data-analysis methods. This not only allowed us to identify the current state of the art with respect to phylogenetic networks, but it also allowed computationalists to improve their understanding of biological data and how biologists proceed to analyze it, as well as allowing biologists to obtain immediate feedback with respect to their data-analysis issues.

Production of phylogenetic networks seems to have come a long way in the past few years, although there is still no single "one-stop shopping" software tool to use. Practical issues getting programs to perform on all computer types were identified, along with data-format issues. Nevertheless, all of the participants seemed to find that this was a very valuable exercise, as a means of focussing interactions among themselves.

Finally, we considered both European and U.S. funding for network research, in the latter case assisted by David Mindell (from the N.S.F.). In particular, we identified sources of funding for future workshops (either in the south of France or the north-eastern U.S.A.).

The canal-boat cruise turned out well, in spite of the somewhat uncooperative weather. The football, of course, has turned out to be rather disappointing for the hosts, although they have one more game to play.

Tuesday, July 8, 2014

Touching the Data, report 1


We have now completed two days of the workshop. We have had a relaxed approach to progress, and are thus currently running behind the nominal schedule. Nevertheless, we are progressing splendidly.

We had three talks on the first day and one today. I tried to kick things off by asking a series of what I consider to be unanswered questions from observing practitioners and computationalists in action, although apparently several members of the audience already had their own answers to some of these. The bottom line is that phylogenetic analysis focuses on data patterns while interpretation focuses on processes / mechanisms, and this constitutes a large part of the apparent separation of practitioners and computationalists.

Steven Kelk and Luay Nakhleh introduced the diversity of computational approaches that we already have. These presentations neatly complemented each other, providing a valuable summary of the field as well as an overview of current limitations and future prospects. This topic was taken up later by various members of the audience, as one of the inherent problems for practitioners is how to navigate through the methods to choose a suitable one -- there are methods based on parsimony, likelihood and bayesian analysis, and methods that tackle de novo network construction, gene tree / species tree reconciliation, gene tree scoring, and network presentation.

This topic was followed up today by presentations introducing some of the currently available software. Some of these have progressed significantly in recent years, notably PhyloNet and Dendroscope, and there are some relatively new ones, as well as even newer ones in the pipeline. Based on the literature, these programs are being dramatically under-used compared to their actual usefulness.

This morning Scot Kelchner introduced us to the application of Zen Buddhism to science in general and phylogenetics in particular. This went down much better than he seemed to be expecting -- there were apparently a lot of  "Zen" people in the room. The basic idea is not to get trapped by preconceived expectations, especially arbitrary categorical notions, when interpreting the output of a phylogenetic analysis. You can consult The Nine-Headed Dragon River, by Peter Matthiessen, if you would like further information.

Finally, we got to the topic implied by the workshop's title: Touching the Data. We had a brief run-through of the pre-existing datasets stored with this blog (see the upper right-hand corner), which cover some of the diversity of what practitioners have provided to date in the way of usable datasets with "known" phylogenetic patterns.

By far the most interesting, however, was the presentation of some recent datasets made available by members of the workshop, notably Axel Janke (bear species), Scot Kelcher (bamboo species) and Mattis List (Indo-European languages) (Jim Whitfield will present his datasets tomorrow morning). These datasets generated much interest, as they provide a diversity of different possible applications for phylogenetic networks. The idea from here on in the workshop is to address what can currently be done with these datasets and what we might like to do with them if the tools were available. This will help focus the participants on specific practical issues, which should lead to the progress that we hope to achieve.

It has rained most of the day, which is actually unusual -- intermittent rain is more common in this climate. We are currently waiting for the football to start: Germany versus Brazil. Tomorrow will be the Netherlands versus Argentina. It is risky being in this country this week! The current local betting is for an all-European final,an assessment that involves no cultural bias whatsoever.

Wednesday, July 31, 2013

Trends in Genetics: The Future of Phylogenetic Networks


A couple of weeks ago I reported on those journal covers that I know illustrate phylogenetic networks. I am happy to report that networks have now also made it onto the cover of Volume 29 Issue 8 of Trends in Genetics. The cover illustration combines the traditional tree metaphor for phylogenetics with the new metaphor of a network.


The cover story is the review article by Eric Bapteste, Leo van Iersel, Axel Janke, Scot Kelchner, Steven Kelk, James McInerney, David Morrison, Luay Nakhleh, Mike Steel, Leen Stougie and James Whitfield: Networks: expanding evolutionary thinking, on pages 439-441.

The article is one of the tangible outcomes of the workshop last October, at the Lorentz Center in The Netherlands: The Future of Phylogenetic Networks. The workshop participants agreed that we should be active in promoting the use of networks for evolutionary analyses, and this article, written by a group of biologists and computational biologists, seeks to do just that.

There will be further outcomes of the workshop, including follow-up meetings at the same venue.

Saturday, October 20, 2012

The Future of Phylogenetic Networks: Photos

Evidence that we were in the Netherlands.



The Lorentz Center building. The Center offices are to the right and left of the yellow window.



The seminar room. Left to right: Axel Janke, David Morrison, Hans-Jürgen Bandelt, Steven Kelk, Mike Steel.


The dinner cruise around the canals.


Dan Gusfield (obscured at the left) telling the young people how it is. Left to right: Mareike Fischer (plus dog), Chris Whidden, Leo van Iersel, Steven Kelk, Simone Linz, and Céline Scornavacca (resting).


Left to right: James Oldman (chopped off), Juan-Diego Santillana-Ortiz (green jacket), Irma Lozada Chávez (at back), Jack Koolen, Axel Janke, David Morrison, Magnus Bordewich (back to camera), James McInerney, Mike Steel (obscured), and Charles Semple.


Three men and a goat, at Leiden market. Left to right: David Morrison, Scot Kelchner, and Jim Whitfield.


Friday, October 19, 2012

The Future of Phylogenetic Networks: Day 5


There were only two talks today. First, Leo van Iersel provided an excellent summary of the week's talks and discussions. He put all of the points into context that had been made throughout the week, and provided a clear context for future activities.

Finally, Daniel Huson provided the keynote talk for the week. He especially covered the need for algorithms to address real data (eg. nonbinary trees, multiple trees, partial trees), for the output to return all relevant solutions, and to provide tools for visualization and interactive analysis.

All in all, a very productive week was had by all. The speakers are especially to be thanked, for providing overviews of and introductions to the various topics; as well as those people who contributed the most to the discussion sessions. The communication between the biologists and the mathematicians was excellent (they are, after all, simply two parts of a continuum), with everyone learning something new as well as making new friendships. Hopefully, this workshop will be the start of a series of such meetings, which would then be a forum for building a substantial body of work on the topic of phylogenetic networks.

The Future of Phylogenetic Networks: Day 4


There were two talks today and two lengthy discussion sessions.

Hans-Jurgen Badlet started the day with a smallish audience that increased as time progressed. This may have something to do with the noise emenating from the hotel bar the previous night. Hans surveyed the field of splits networks, and especially their uses for data quality control in forensic and medical databases. The extent of this use was news to most of the audience.

The morning discussion turned out to be about the extent to which it might or might not be desirable to have some sort of "standards" or even "protocols" for effective use of network techniques. Clearly, many if not most users of phylogenetics are not experts, and misuse or misunderstanding of networks is a real possibility. There was no particular consensus on this issue.

The afternoon discussions covered three topics. First, we discussed ways of detecting hybridization using networks, as opposed to detecting them using directly biological techniques. There were arguments both for and against the successful use of networks, with the consensus being that networks have a useful role.

We then proceeded to a consideration of the extent to which network methodology is prepared for the expected influx of genome-scale datasets. The answer seems to be "not yet", but even with the available methods there is much scope for effective analysis.

The third topic was the practical use of current methods and programs for exploratory analysis of multi-gene data. A number of additions or modifications to current implementations were suggested, including some measure of "tree-likeness" to rank different genes in terms of their data patterns.

The day finished with Charles Semple, who is always first to breakfast and last to leave. He explained the various measures optimizing the process of combining trees, notably reticulation number and its various definitions. There was much use of the whiteboard, which is the mark of a true mathematician.

Wednesday, October 17, 2012

The Future of Phylogenetic Networks: Day 3


There were four talks today and one discussion session. We also spent the evening on a boat cruise around the waterways north of Leiden, supping on an Indonesian buffet and consuming some distinctly non-Indonesian desserts. No-one embarrassed themselves, and so (sadly) there are no stories to be told.

Jim Whitfield started the day's talks by contemplating the varied ways in which entomologists might be interested in using a network, especially with genome-scale data. He eventually worked his way around to the topic of datasets for validating network algorithms, which would need to cover all of these possibilities.

Barbara Gravendeel continued the same theme, by presenting some datasets, mainly involving orchids, for which there is independent evidence that some of the taxa are hybrids. Such datasets could be used for algorithm validation.

The discussion session followed, which focussed on the various ways in which validation datasets could be made available. In the short term, it seems likely that a web page will be set up with this information for each dataset: (i) a link to the online dataset, (ii) a link to the relevant publication(s), and (iii) a brief description of what relevant data patterns are believed to be included in the dataset.

Mike Steel then started the very mathematical afternoon. He mainly contemplated the extent to which non-tree biological processes could create tree-like patterns in the data. There are theoretical ways to differentiate various signatures of gene tree incongruence in the context of triplets, but also sources of inconsistency in phylogeny reconstruction.

Dan Gusfield finished with a coverage of ancestral recombination graphs, including their possible use in biology, but mainly some of the potential things that create problems for reconstruction. The mathematical part of the audience looked very enthused by the end of the afternoon.

The Future of Phylogenetic Networks: Day 2


There were five talks today and one discussion session. We were rained upon for much of the day, although it was better in the afternoon.

Scot Kelchner started by explaining how networks are currently used in botany. Plants have long been accepted as having a complex history, and so there is great potential for using networks to explore this complexity, from detecting unexpected hybrids to assessing lineage sorting. He emphasized the valuable role of networks as a paradigm for both exploratory data analysis and hypothesis generation.

Teun Boekhout addressed the enormous network complexity of fungi, although it seems that very few practitioners are using the available mathematical methods.

James McInerney, somewhat the worse for wear, then discussed microbiology, where much attention has been given recently to horizontal evolutionary processes. He explained the continued dominance of the tree paradigm within prokaryote studies as resulting from interest in the transmission tree of the pathogens and its obvious connection to a phylogenetic tree. He then turned to the Tree of Life, which bacteriologists seem to see as their special preserve, pointing out its inadequacy as a model, before turning to the range of microbiological questions that exist only in the context of a network paradigm.

Eric Bapteste then further explored the concept of network thinking as opposed to tree thinking, emphasizing how limited the tree paradigm is when studying the known diversity of biological phenomena. Most interestingly, however, he also noted that even the network paradigm has limitations for studying biodiversity, as they need to be linked to other types of biological networks.

Vincent Moulton produced the only mathematical talk of the day, covering the mathematical quantification of branch support within networks, which is clearly of interest for both data exploration and hypothesis generation. To date, bootstraps are the only method implemented in the software, although they are rarely used in practice. Delta plots have also been proposed, and are quick to calculate, but there are other theoretical possibilities to be explored.

The topic for the discussion was not pre-determined, and turned out to be just how much automation would be useful in network analyses. The consensus is that an attempt at complete automation would be counter-productive. However, the most thorny issue of debate was exactly which types of network are likely to be most useful to biologists. The problem here is that the mathematicians need an explicit description of such networks in order to produce them, while the biologists do not yet have such a description. The issue remained unresolved when we adjourned to the coffee room.

Tuesday, October 16, 2012

The Future of Phylogenetic Networks: Day 1


There were three talks today and two discussion sessions.

Steven Kelk and I got things rolling by introducing the topic of networks from the mathematical and biological perspectives, respectively. I thought that we both did a good job, but I learned more from Steven's talk than from my own.

Luay Nakhleh then presented some of the computational challenges of moving from a tree perspective to a network. Perhaps the most interesting of these to a biologist is the decreasing independence of reticulations as gene sampling increases (eg. due to gene linkage). Independence is a basic assumption of most computational methods in biology, and the consequences of violating this assumption are rarely addressed. However, for network construction the potential non-independence of reticulations seems to be of fundamental importance for any biological interpretation of reticulation causes.

Computationally, the obvious challenge is the complexity of scoring a network compared to a tree. Calculating the parsimony score of a network is hard enough (although trivial for a tree), and scoring the likelihood is even worse. This calls into question the practicality of using likelihood in the context of networks.

However, perhaps the most interesting challenge is how to model inter-locus incompatibility. Within-locus mutations are currently addressed using substitution/indel models in phylogenetic tree-building, but the special focus of networks is on the inter-locus patterns, about which we know much less in terms of appropriate modelling.

Axel Janke and Katharina Huber finished the day by leading discussions on why so few people currently use networks in phylogenetics, and the obvious cultural divide between mathematicians and biologists, respectively. The biologists seemed to dominate the first discussion and the mathematicians the second one.
In the former case, the main conclusion from the discussion was that the current phylogenetic "culture" is focussed so strongly on trees that the extra benefit of using a network is not obvious to practitioners. Indeed, there is still considerable focus on getting people to think in terms of trees rather than linear evolution, so that moving on to the complexity of networks simply confounds the situation. Suggestions were forthcoming about how we could be proactive in addressing this issue, including increasing the profile of networks in journals, but also the need to provide more biological information from analyses than merely the network topology.

As for the cultural divide, this is a long-standing issue that arises from the different thought processes involved in mathematics and empirical science, and the consequent differences in language. The consensus was that there are no hurdles that can't be overcome given sufficient time and patience. Moreover, trans-disciplinary people are becoming more common, which nullifies many of the potential problems.
So, a productive start to the workshop was made, which bodes well for the rest of the week. Sadly, this was the sunniest day since I arrived in the Netherlands, and I spent it indoors!