Showing posts with label Placentals. Show all posts
Showing posts with label Placentals. Show all posts

Wednesday, June 8, 2016

Why do so few biologists look at their phylogenetic data?


Most data analyses involve processing the data using some model. For example, standard parametric statistical tests assume a normal distribution for the "error" term, as well as equal variances and linear relationships between the variables. If these model assumptions ado not hold, then any inferences from the tests may be incorrect.

It is possible to look at any dataset in a model-free manner, although this does not necessarily lead to any strong inferences. Looking at data is usually called exploratory data analysis. This is often done using graphs of various types.

Exactly the same principle applies to phylogenetics. A phylogenetic tree is an inference from the data via a given model. The inference is a reconstructed genealogical history assuming a divergent tree. In this context, different models will often (usually?) give different inferences.

Therefore, most phylogeneticists never actually see their data. What they see, instead, is the data as processed through some model. That is, they see inferences from the model, not the original data. Models are important, but the data should be even more important, for a scientist.

It is thus interesting that so many phylogeneticists skip the step of looking at their data, and proceed immediately to the model-based inference. So many of the disagreements throughout the literature end up being about the models and not the data. There are very strong opinions about which models should be used, with less attention being paid to whether the data contain sufficient information to answer the original scientific question in the first place.

A specific example of this was discussed in some earlier blog posts:
Conflicting placental roots: network or tree?
Why are there conflicting placental roots?
In this example there are three possible genealogical patterns, each of which has been reported to receive strong support from model-based tree inference of nucleotide sequences. However, when looking at the sequence data themselves, in a model-free manner using data-display networks, any one dataset shows all three possible patterns. So, any inference of a single tree is coming from the model not from the data. That is, the data do not distinguish between the three genealogies, but the models do discriminate amongst them.

It is worth mentioning here that a haplotype network is not a genealogy. Instead, it is a summary of a population dataset, which may contain some phylogenetic patterns or it may not. So, a haplotype network is closer to exploratory data analysis than it is to model-based inference. This point is clearly made by Jessica W. Leigh and David Bryant (2015. PopART: full-feature software for haplotype network construction. Methods in Ecology and Evolution 6: 1110-1116):
The haplotype networks do provide, however, a concise and accessible representation of the data themselves, one aspect which is often lost in methods heavily dependent on model-based inference.
Looking at the data before you start processing it can be a very good idea. After all, you may be able to avoid unlikely inferences.

Wednesday, August 28, 2013

Why are there conflicting placental roots?


Last week I noted that there has been recent activity concerning the "placental root" problem, in which different genetic datasets support different phylogenetic trees for the root of the placental mammal clade (Conflicting placental roots: network or tree?). There are two articles (by Morgan et al. and Romiguier et al.) in the current issue of Molecular Biology & Evolution that address this problem with genomic data, and find two different well-supported trees.

This is an issue that I also addressed in a much earlier post (EDA or post-optimality analysis of phylogenetic data?), based on the genomic dataset of Meredith et al., in which I concluded:
It is not immediately obvious that a tree-building analysis is going to be of much use for this dataset. There is certainly some "power of building phylogenies from large densely sampled datasets", but this does not automatically mean that those phylogenies will be tree-like. Evolution involves a more diverse process than that.
In all of these cases, sophisticated substitution models (nucleotide or amino acid) were used as the basis for building a phylogenetic tree, whereas the network analysis of Hallström & Janke suggests that mammalian evolution may not be strictly bifurcating.

My interest in this blog post is in investigating the relative roles on the data and the substitution models in producing the phylogenetic trees. I use splits graphs of the recent data (using the SplitsTree program) as an exploratory data analysis, to visualize the signals in the datasets and which trees they might support under different circumstances.

The analyses

Any phylogenetic analysis depends on the quality of the data, in terms of the sampling of both taxa and characters. Both Morgan et al. and Romiguier et al. used the protein-coding sequences for most of the 40 currently available mammalian genomes.

However, it is worth noting at the outset that the sampling of the root taxa is rather poor. The root involves the relative relationships of the Xenarthra and the Afrotheria, and yet there are only two sampled Xenarthra species and three sampled Afrotheria (the remaining taxa are split between the Laurasiatheria and Euarchontoglires). Perhaps we are asking too much in expecting these data to resolve the root at all.

We can start the investigation with the data of Morgan et al., based on the concatenated amino acid sequences. The first NeighborNet analysis uses the simplest model possible, the hamming distance (which is simply the number of alignment differences between the taxa). I have colour-coded the four taxonomic groups, for convenience.


Note that all four taxaonomic groups appear to be monophyletic (ie. they are each supported by a unique split), as also is the Xenarthra+Afrotheria group. However, the raw data attach the outgroup to the placental group away from both the Xenarthra and the Afrotheria. Indeed, the data suggest that the Insectivora (Sorex+Erinaceus) are candidates as the sister to the rest of the placentals.

The effect of the substitution model on the data analysis can be evaluated by including a more sophisticated genetic distance. I have chosen the JTT amino-acid model, with the inclusion of a proportion of invariant sites (estimated by SplitsTree to be 30%). The corresponding NeighborNet is shown in the second graph.


This network attaches the outgroup near the "expected" taxa (Xenarthra, Afrotheria), although the location of Sorex is rather problematic. However, the split supporting the group Xenarthra+Afrotheria as the sister to the rest of the placentals is still very small, being ranked only 28th of the 82 non-trivial splits that involve at least one placental species. So, even this simple model does not provide strong support for the root location. However, it seems obvious that the root location is being determined as much by the substitution model as by the data, suggesting that the data cannot provide convincing evidence alone.

We can now proceed to study the data of Romiguier et al., based on the maximum-likelihood gene trees (GTR+GAMMA model) from the 560 genes, rather than the original alignment data. Here I have used a Consensus Network that displays all of those splits occurring in at least 24% of the trees. This percentage is the smallest that produces only a single reticulation in the network.


So, the most ambiguous part of the set of trees (ie. where there is most conflict among the trees) turns out to be where the outgroup attaches to the placental group. This is hardly surprising. What is more interesting is that the split support for each of the three alternative attachment points is very similar:
Outgroup+Xenarthra  0.00566
Xenarthra+Afrotheria 0.00557
Outgroup+Afrotheria  0.00496
So, the gene-tree data do not favour any one of the three alternative placental roots.

Conclusion

It is clear from these exploratory analyses that the genomic data do not, on their own, provide conclusive evidence regarding the root of the placental clade. The approach of Morgan et al. and Romiguier et al. has been to use a tree model based on sophisticated substitution models, thus arriving at conclusions that depend as much on their models as on the data. They used different models and got different trees, based on roughly the same data.

This is one approach to phylogenetics, to use more sophisticated models; but an alternative is to recognize that evolution itself is sophisticated, and therefore does not necessarily produce a dichotomous tree. In this case, it seems more likely that the conflicting signals at the placental root reflect non-tree-like processes (such as hybridization), so that tree-based analyses are inappropriate, no matter how fancy the models are.

References

Hallström, Janke (2010) Mammalian evolution may not be strictly bifurcating. Molecular Biology & Evolution 27: 2804-2816.

Meredith et al. (2011) Impacts of the Cretaceous terrestrial revolution and KPg extinction on mammal diversification. Science 334(6055): 521-524.

Morgan et al. (2013) Heterogeneous models place the root of the placental mammal phylogeny. Molecular Biology & Evolution 30: 2145-2156.

Romiguier et al. (2013) Less is more in mammalian phylogenomics: AT-rich genes minimize tree conflicts and unravel the root of placental mammals. Molecular Biology & Evolution 30: 2134-2144.

Wednesday, August 21, 2013

Conflicting placental roots: network or tree?


In this blog we champion networks as a fundamental model for phylogenetics. Networks are more general than trees, in the sense that some networks are more tree-like than are others. However, I have noted before that the current trend in phylogenetics seems to be to try to use more and more complex trees as the phylogenetic model, rather than embracing networks as a more flexible model (Resistance to network thinking).

An interesting example of this trend is in the current issue of Molecular Biology & Evolution. There are two articles that investigate the root of the placental clade, by Morgan et al. and Romiguier et al., along with an editorial commentary by Teeling & Hedges.

The "placental root" problem has been difficult to resolve as a bifurcating process because different genetic datasets support different trees. As noted by Teeling & Hedges: "Untangling the root of the evolutionary tree of placental mammals has been nearly an impossible task. The good news is that only three possibilities are seriously considered ... Now, two groups of researchers have scrutinized the largest available genomic data sets bearing on the question and have come to opposite conclusions". The three alternative tree histories for the clade root are shown in the figure.


Both of the new empirical studies are based on the protein-coding sequences for most of the 40 currently available mammalian genomes. Morgan et al. use heterogenous substitution models to account for tree and dataset heterogeneity, and get strong support for option (c). Romiguier et al. divide their dataset into GC-rich and AT-rich genes, conclude that the GC-rich genes are most likely to suffer from long-branch attraction, and get strong support from the AT-rich genes for option (a).

Teeling & Hedges continue: "Needless to say, more research is needed." No! Previous genome-scale analyses of more than one million amino acid sites from orthologous protein-coding genes have not rejected any of the three alternatives, despite the statistical estimate that 20,000 amino acid sites should be sufficient to resolve the question at this level of divergence given the tree structure, branch lengths, and number of substitutions (Hallström & Janke 2010). Doesn't this mean that we have enough evidence already?

Clearly, the conflicting results should lead the reader to at least consider the idea that something might be wrong with the underlying tree model itself. Both of these new analyses are still based on tree models, no matter how sophisticated those models might be (see also the several other papers cited by Teeling & Hedges), and no matter how much data are involved.

An alternative perspective is provided by Hallström & Janke (2010): "Mammalian evolution may not be strictly bifurcating". Their network analysis of retroposon insertion data supports an alternative hypothesis for the history of placentals: the early divergences involved incomplete lineage sorting and hybridization. Neither of these two evolutionary processes is accounted for in the tree models of Morgan et al. and Romiguier et al., but both can be integral parts of a network model.

Conclusion

I think that we can see the suggested move from trees to networks as a form of Kuhnian paradigm shift. In Kuhn's historical model, during the period of "normal science" the failure of results to conform to the current paradigm is not seen as refuting the paradigm, but instead is seen as resulting from errors by researchers (e.g. use of inadequate models, acquisition of unreliable data). However, in the Kuhn model, as anomalous results accumulate a new paradigm emerges that subsumes the old results along with the anomalous results, forming a single new framework or paradigm.

Non-tree-like phylogenetic results are currently not seen by most phylogeneticists as refuting the paradigm of a phylogenetic tree, but instead are the result of inadequate phylogenetic tree-models and/or insufficient data (as exemplified by Salichos and Rokas 2013). Nevertheless, these results can also be seen as refuting that paradigm. In that case, a shift to network thinking would embrace all of the tree results as well as the non-tree ones, and would thus form a viable new paradigm.

We should not really call this a Kuhnian "revolution", of course, since tree-thinking and network-thinking are not incompatible, but rather the one is an extension of the other.

Note: There is a follow-up post — Why are there conflicting placental roots?

References

Hallström BM, Janke A (2010) Mammalian evolution may not be strictly bifurcating. Molecular Biology & Evolution 27: 2804-2816.

Morgan CC, Foster PG, Webb AE, Pisani D, McInerney JO, O’Connell MJ (2013) Heterogeneous models place the root of the placental mammal phylogeny. Molecular Biology & Evolution 30: 2145-2156.

Romiguier J, Ranwez V, Delsuc F, Galtier N, Douzery EJP (2013) Less is more in mammalian phylogenomics: AT-rich genes minimize tree conflicts and unravel the root of placental mammals. Molecular Biology & Evolution 30: 2134-2144.

Salichos L, Rokas A (2013) Inferring ancient divergences requires genes with strong phylogenetic signals. Nature 497: 327-331.

Teeling EC, Hedges SB (2013) Making the impossible possible: rooting the tree of placental mammals. Molecular Biology & Evolution 30: 1999-2000.