Showing posts with label Phylogenetic tree. Show all posts
Showing posts with label Phylogenetic tree. Show all posts

Monday, June 1, 2020

To what degree are Median-joining networks phylogenetic?


In a comment to the recent paper by Forster et al. (2020), Sánchez-Pacheco et al. (2020) argue that Forster et al.'s analysis is "neither phylogenetic nor evolutionary" because it's based on the use of a Median-joining network. They don't re-analyse the data, but instead mostly refer to a paper they published four years ago in Cladistics (Kong et al. 2016), the journal of the Willi-Hennig Society.

In that earlier paper, Kong et al. conclude:
Other than fast computation and very attractive graphics, MJNs [Median-joining networks] harbour no virtue for phylogenetic inference. MJNs are distance-based, unrooted branching diagrams with cycles that say nothing about the evolutionary history due to the absence of direction. MJ was introduced in 1999 and, in contrast to most scientific ideas, its application has spread rapidly through copying the methods of others, and, unfortunately, with little further scrutiny. We hope that the theoretical arguments presented here can reverse this trend.
It seems unlikely that it will, as I will argue here.

What makes a graph a phylogenetic tree or network – direction

Kong et al. argue that a line graph needs to be directed (ie. the edges indicate a time direction) in order to represent a phylogeny, which is a good point. After all, a phylogenetic tree is a directed (rooted) branching diagram that represents the hypothesized relationships among the organisms under study.

A phylogenetic tree (see also: Fritz Müller and the first phylogenetic tree)

A phylogenetic network is the generalization of a phylogenetic tree, as it combines lineage splits (divergences) with lineage anastomoses.

A phylogenetic network including a reticulation leading to a circle in the graph — B is the product of crossing of lineages that produced its sisters A and C.

Since a MJ network is, per se, an undirected graph, it thus cannot be an explicit phylogenetic network.

However, following this argument, few inferred trees are directly a phylogenetic tree, either — including the Nextstrain-generated tree on the GISAID page that is promoted by Mavian et al. 2020 (which is another comment to Forster et al., focusing on data issues). Irrespective of which criterion we use to optimize the tree, almost all trees we infer (with no matter what tree inference software) are unrooted graphs — in general, we root them only after the analysis, by defining one leaf or a subtree as an outgroup. (Note, this includes those based on parsimony, the method of choice of the Willi-Hennig Society and Cladistics to this day.)

The difference between inference and interpretation: Using the tip sequences, we can infer a single most parsimonious (6-step long; using PAUP*'s branch-and-bound or NETWORK's MJ algoritm), but also most likely and shortest (distance-based), unrooted tree. By defining a root – here: one taxon designated as outgroup and assuming that all single-taxon-unique sequence patterns are autapomorphies – we can interpret the inferred tree as four different phylogenetic trees.

The same can be said of MJ networks — outgroup rooting can be applied (Finding the CoV-2 root).

Difficulty in depicting ancestor-descendant relationships

A phylogenetic relationship focuses on ancestors, which, for the purpose of inferring a phylogenetic tree, are considered to be purely hypothetical, although they are not hypothetical in a MJ network (or related graphs). We can easily create character sets where the inferred tree will not "represent the hypothesized relationship". Most parsimony studies show a strict consensus cladogram of most-parsimonious trees (MPT). This is unproblematic, as long as all leaves have the same age, and all of the cladogenic events resulted in unique, lineage-conserved character patterns. We then:
  1. would only infer but a single MPT;
  2. have no zero-length branches.
So, following Hong et al.'s logic, any dataset that results in more than one MPT and has subtrees including zero-length branches (like our example above) cannot qualify as phylogenetic trees.

Median-joining networks are, like MP trees (both use parsimony as the optimality criterion), vulnerable to homoplasy (Using Median networks to study SARS-CoV-2; see also Mavian et al. 2020), but while a MP tree (or any other tree we infer) cannot resolve ancestor-descendant relationships, MJ networks can (see eg. Why do we still use tree for Neanderthal genealogy).

Median (or MJ) network, left, and MPT, right, inferred from a perfect matrix. "x" = all-ancestor, ie. represents the root. "a" is the ancestor of "B" and "C", "d" of "f" to "H", "f" of "G" and "H". While the median networks depicts all ancestor-descendant relationships, the MPT only depicts them indirectly by trichotomies including the ancestor as zero-length branch.
Imperfect matrices (data including homoplasy) lead to wrong edges and branches. Being able to recognize ancestors, the MJ network comes closer to the phylgoenetic tree (same as above; from Clades, cladograms, cladistics, and why networks are inevitable).

Hence, Bandelt et al.'s (1999) statement, as cited by Kong et al., that “reconstructing phylogenies from intraspecific data ... is often a challenging task because of large sample sizes and small genetic distances between individuals”. Such data results in largely uninformative, comb-like MPT strict consensus trees. This is because identical sequences, equally probable alternative pathways, non-dichotomous differentiation patterns, and ancestral sequence variants present in the data increase the number of MPTs (sometimes to near-infinity). This leads to the collapse of branches in the strict consensus tree used to summarize the MPT sample. Probabilistic methods struggle, too, because the likelihood surface of the tree space is too flat to make a call.

[Kong et al. point to the mathematical definition of 'network', as "nothing more than an unrooted branching diagram with reticulation" but not of 'tree', which they consider is always a directed acyclic graph, ie. synonym to 'phylogenetic tree'. However, it is, inference-wise, clearly nothing more than an unrooted branching diagram without reticulation.]

Confusing heuristics with principle

To discredit the MJ network, Kong et al. then "... focus on its phenetic nature."

There is a tendency among cladists to dismiss a method as "distance-based", as this is treated as synonymous with phenetics. In reply, Joe Felsenstein commented on this alleged fundamental difference between distance-based and parsimony methods of tree inference (Felsenstein 2004, Chapter 10, p. 145f, The irrelevance of classification):
The terminology is also affected by the lingering emphasis on classification. Many systematists believe that it is important to label certain methods (primarily parsimony methods) as "cladistic" and others (distance matrix methods, for example) as "phenetic". These are terms that have rather straightforward meaning when applied to methods of classification. But are they appropriate for methods of inferring phylogenies? I don't think they are. Making this distinction implies that something fundamental is missing from the "phenetic" methods, that they are ignoring information that the "cladistic" methods do not. In fact, both methods can be considered to be statistical methods, making their estimates in slightly different ways.
The following chapter in Felsenstein's book (Chapter 11, pp. 147–175) deals exclusively with the "phenetic" distance matrix methods because they were the first to be used to infer phylogenetic trees (their limitations are outlined on pp. 174f).

Because the inference of MJ network starts from the generation of a Minimum-spanning network, which is generated from a distance matrix, Kong et al. argue the MJ network is merely a distance-based graph, ie. "phenetic", and "not phylogenetic". Any NP hard problem requires heuristics but, just because we use a distance-based graph to start with, doesn't determine whether the end-product is or is not a distance-based graph.

For instance, the Neighbor-joining (NJ) algorithm (Felsenstein 2004, p.166ff) is a cluster algorithm, which finds a phylogenetic tree fulfilling either the Minimum evolution (ME, p. 159f) or Least-squares (LS) criteria (p. 148ff). Thus, the tree inferred is, indeed, based on a distance-matrix via NJ, but it is not a cluster dendrogram — instead, it is a ME or LS optimised phylogenetic tree. Similarly, FastTree, IQTree, and RAxML are extremely fast programmes to infer Maximum likelihood (ML) phylogenetic trees; but, while FastTree and IQTree start with "phenetic" Neighbor-joining trees, RAxML (like GARLI before) infers first a quick-and-dirty parsimony tree. The final product in both cases is a topology optimized under ML, and the results are hence ML trees and not distance-based or MP trees (even though they started that way).

The final MJ network shows the most parsimonious evolutionary pathways that change one sequence type into another. When you infer it with the NETWORK program, all inferred mutations are mapped onto the final graph, and, using the Steiner post-analysis step, you can look through all of the MP trees that have been included in this graph. However, according to Kong et al. these are not MP trees:
[Following Farris (1970] Invoking principles of parsimony does not validate a phenetic technique as being a phylogenetic method. Indeed, the best Steiner trees are not necessarily the most parsimonious trees.
Kong et al. did not provide any real-world data examples; possibly because they would be very difficult to find. Just take my simple example above — clearly the MJ tree is actually a most-parsimonious solution to the data. Alternatively, you could take any data set for which you can infer plausible MPTs with (ie. data where the rate of change is low), eg. using the TNT program, and compare the result – the Consensus network of all MPTs, not the collapsed strict consensus tree (Stop using cladograms!) – with the Steiner trees inferred using NETWORK and the MJ algorithm.

Are medians ancestors? Do cycles represent reticulation?

Kong et al.'s final point is:
BEA99 [ie. Bandelt et al. 1999] stressed that median vectors can be interpreted biologically as existing unsampled or extinct ancestral sequences (i.e. they can represent missing intermediates; Fig. 3). However, a median vector in an MJ analysis is a sequence generated by majority, and is a mathematically drawn point in the final MJN that connects a triplet of sequences. The resulting “evolutionary paths in the form of cycles” (BEA99, p. 37) merely illustrates the failure of the algorithm to choose between alternative, equally optimal connections due to the modification of Kruskal’s algorithm. Consequently, a cycle represents an analytical artefact rather than an evolutionary scenario (Salzburger et al., 2011).
It is obvious to anyone who has ever used MJ networks, and is familiar with their own data, that Medians are likely to be ancestors, and that medians separated by parallel edge bundles are usually alternative ancestors. But, like all inferences, MJ networks may not capture the complex truth.

A phylogeny involving a recombination ...
... and the MJ network that can be inferred on the same data including two wrong edges (red). The West-1/East-ancestor recombinant is resolved as the product of hybridisation of the West- and East-ancestors, while West-2, a descendant of West-1, is resolved as hybrid of West-1 and the recombinant. Any tree included in the network would have 7 steps (ie. is most-parsimonious).
These reconstructed medians thus do bring Kong et al. to their only valid point, which, however, doesn't apply to the method as proposed, but is instead a common misinterpretation of MJ networks — their cycles do not necessarily reflect reticulation.

Bandelt et al. (1999) clearly state that the MJ network is only an approximation, to deal with complex situations. The cycles usually represent equally optimal alternative pathways, and are usually the result of homoplasy but not reticulation. The final goal is hence to get a graph with as few reticulation as possible but as many as are necessary (see NETWORK manual on selection of the epsilon parameter and weighting).

The Sanchéz-Pacheco et al. critique of Forster et al.

As I showed in an earlier post using actual CoV-2 data, only this part of anchéz-Pacheco et al.'s critique of Forster et al.'s paper is valid — we do need to be very careful before we interpret parallel edge bundles in virus-based (or other) MJ networks as being evidence for reticulation. MJ networks can be phylogenetic networks, but they are still consensus networks of competing, equally parsimonious alternatives. If we take a strict position, then most MJ networks are probably not phylogenetic networks; but neither are all trees phylogenetic trees.

Everything else in their comment is simply cladistic lore. Most importantly, their critique ignores the fact that the obvious alternative to MJ networks when analysing low-divergent virus data, which is parsimony-based trees, has exactly the same data-inherent shortcomings — ie. vulnerability to homoplasy, impossibility to detect and reconstruct recombination. They also have an extra one: they treat all samples to be of the same age and generation, and thus have to resolve actual ancestors as being sisters. Which increases the number of possible, equally parsimonious solutions.

The 19, 7-step long MPTs that can be inferred for the recombination example using PAUP*'s branch-and-bound algorithm – rooted with the Source, the common ancestor of all ("AllAnc") – and their strict consensus tree (gray background, 11-steps long). "Best" shows a phylogenetic tree that comes closest to the true tree: ancestors are resolved as zero-length tips in clades including their descendants. "Close" denotes trees that only misplace the recombinant (purple), which – being a recombinant of the East ancestor (red) and West-1 (blue) – should be placed in a tree as sister to either parent.

The consequence of this is that what Sánchez-Pacheco et al. and Kong et al. criticize about the MJ networks applies even more to the predominately used phylogenetic trees. As David pointed out earlier (Problems with the phylogeny of coronaviruses): virus trees may be inferred using phylogenetic methods but they effectively depict only similarity patterns.

References

Bandelt H-J, Forster P, Röhl A. 1999. Median-joining networks for inferring intraspecific phylogenies. Molecular Biology and Evolution 16:37-48.

Forster P, Forster L, Renfrew C, Forster M. 2020. Phylogenetic network analysis of SARS-CoV-2 genomes. PNAS 117:9241–9243.

Felsenstein J. 2004. Inferring Phylogenies. Sunderland, MA, U.S.A.: Sinauer Associates Inc.

Mavian C, Kosakovsky Pond S, Marini S, Rife Magalis B, Vandamme A-M, Dellicour S, Scarpino SV, Houldcroft C, Villabona-Arenas J, Paisie TK, Trovão NS, Boucher C, Zhang Y, Scheuermann RH, Gascuel O, Tsan-Yuk Lam T, Suchard MA, Abecasis A, Wilkinson E, de Oliveira T, Bento AI, Schmidt HA, Martin D, Hadfield J, Faria N, Grubaugh ND, Neher RA, Baele G, Lemey P, Stadler T, Albert J, Crandall KA, Leitner T, Stamatakis A, Prosperi M, Salemi M. 2020. Sampling bias and incorrect rooting make phylogenetic network tracing of SARS-COV-2 infections unreliable. PNAS doi:10.1073/pnas.2007295117

Sánchez-Pacheco S, Kong S, Pulido-Santacruz P, Murphy RW, Kubatko L. 2020. Median-joining network analysis of SARS-CoV-2 genomes is neither phylogenetic nor evolutionary. PNAS, doi:10.1073/pnas.2007062117.

Monday, October 21, 2019

Why the emperor has no clothes on – the mighty matK


In a recent paper published in PeerJ, Walker et al. (2019) take a close look at the complete plastome data of angiosperms. Although they don't find anything fundamentally new — well, at least not for those of us who have looked at the oligogene datasets we worked with — it's nice to see that somebody has been willing to do it in a very comprehensive way, and thereby published what some of us have long known:
  • A combined tree is not the sum of the genes that have been combined;
  • Single-gene trees can tell you very different stories.
Even if the overall branch support is pretty high, we always should be aware of internal data conflict.

When looked at closely, the emperor, in this case the Angiosperm Phylogeny Group (APG) complete plastome tree, maybe not be entirely naked, but is clothed in very few of the many garments at his disposal. Effectively the branches in the plastome reference tree draw their support from very few of the 79 genes/gene regions in the plastome.As Walker et al. note:
"Of the most commonly used markers, matK, greatly outperforms rbcL; however, the rarely used gene rpoC2 is the top-performing gene in every analysis. We find that rpoC2 reconstructs angiosperm phylogeny as well as the entire concatenated set of protein-coding chloroplast genes."

Fig. 1 from Walker et al. showing the (lack of) individual gene support for the angiosperm reference phylogeny.

However, there is one aspect of the paper that calls for a network-based blog post:
"Following the typical assumptions of chloroplast inheritance [i.e. that the entire plastome shares a common history being passed on solely by the mother in angiosperms], we would expect all genes in the plastomes to share the same evolutionary history. We would also expect all plastid genes to show similar patterns of conflict when compared to non-plastid inferred phylogenies ... Our results, however, discussed below, frequently conflict with these common assumptions about chloroplast inheritance and evolutionary history."
Getting incongruent branches in the single-gene trees, including a few highly supported ones, is taken as evidence for different histories potentially mixed within the plastome. Walker et al. give references for (potential) recombination and reticulation in plastomes.

I asked a question about whether this logic isn't a bit naive about tree inference. In their response, they pointed to the paper by Sullivan et al. (Mol. Biol. Evol. 2017) — these authors made test for recombination in Picea (spruce) plastomes, then split the complete plastomes into three structural units, and found two embedded conflicting phylogenies, as shown in the next figure.

Fig. 4 from Sullivan et al. (2017). F1 and F2 are structural regions comprising most of the large single-copy unit, the F3 the two (duplicate) inverted-repeat regions and the small single-copy unit of the Picea plastomes.

This seems to be a compelling case (but note the BS < 100 for conflicting critical branches). It is also quite possible, since gymnosperm plastomes, in contrast to angiosperms, may be paternally or bipartentally inherited. But, is it a valid assumption that each single-gene tree (or, in Sullivan et al.'s case, trees based on multigene regions) reflects the true tree of that gene or gene complex? That is, even if I assume that all of the genes in my matrix share the same history, must they support the same inferred tree?

Since I have worked a lot at low taxonomic levels, and often with other people's (plastid) data (during my entire career, I remained faithful to the nuclear-encoded ribosomal DNA spacers), my spontaneous answer would be: Absolutely not! Topological conflict may hint towards decoupled gene histories — it is a neccessary criterion but not a sufficient criterion.

There are quite a lot evolutionary scenarios that will lead to data inevitably supporting wrong branches, or false positives (see also Walker et al.'s discussion). Even if evolution is a strictly dichotomous process (which it clearly isn't):
  • low divergence may result in primitive (underived) sequences ('genetic symplesiomorphies') being shared by distant taxa
  • high divergence may result in saturation, which ultimately triggers branching artifacts
  • long isolation coupled with small active population sizes, repeated bottleneck / massive extinction events and/or lack of radiations will lead to sequences that are different from anything else in our data (in angiosperms, this phenomenon has a name: Ceratophyllum).
In fact, the very argument for angiosperm molecular phylogeneticists to move away from using single-gene phylogenies was that these first single-gene trees had branches that made little sense, especially when based on plastid data.

Single-gene trees will get things wrong. The more signal we add, usually by adding additional gene regions, the more we will reduce these errors (this is best-case scenario, but see Delsuc et al., Nature Rev. Genet. 2005). Thus, if some gene-trees conflict more with the combined tree than do others, it can be for two possible reasons:
  1. The conflicting genes had indeed different evolutionary histories. However, this would have to involve intra-plastome recombination and heteroplasmy, which so far have been very rarely documented in angiosperms.
  2. All genes had the same evolutionary history, but some of the data get more aspects of this true tree right than do others (and, of course, some are wrong that others get right).

And the matK said: "I'm your lord, follow my lead"

Walker et al. (all their scripts and results files can be found on github) find that it's only a few of the genes that essentially make up the combined tree. One of them is an old reliable pal of angiosperm phylogeneticists, the chloroplast matK gene. The literature is full of "multigene" trees that are effectively matK gene-trees using enlarged matrices. The matK determines a topology, and by adding genes that cannot compete with it (being too conserved, too variable or just inconsistently different), we re-enforce this topology. Only branches unresolved by matK will be further optimized using the added data.

Let's look at an example.

For the purpose of this post (and the follow-up), I'll use an old angiosperm matrix on stock (I know the quirks of this matrix). For analysis, I eliminated all of the OTUs with missing gene partitions, mainly to make sure that all of the trees and bootstrap (BS) pseudoreplicate trees have the same set of leaves, so I can summarize the tree samples using consensus networks.

Here's the my combined tree, unpartitioned.

Gray – current APG IV classification, "gold tree" (primary relationships within Mesangiospermae still a matter of debate)
And here is the fully partitioned one (over-parametrized; with each gene/codon position treated as data partition).

Essentially the same tree (some branches elongated, others shortened), eudicot clade and the Ceratophyllum-monocot clade swapped positions. Both trees have the same scale.

Even though my matrix includes only relatively few genes (just 21,550 sites), the tree gets the main aspects of the APG IV standard tree. The support for most of the branches is nearly unambiguous (irrespective of data partition), with the exception of some deep-down relationships within the Mesangiospermae (a long-standing issue, called the "dirty dozen"). The fact that the unpartioned and partitioned analysis agree for most part, indicates the signal in my matrix has no model-related issues (at least, none we could fix by using "better" models).

And the matK tree mirrors the fully partitioned tree, as shown here.

A tanglegram of the matK and the combined trees. Shown is the matK BS support for shared and conflicting edges. Orange asterisks, the monocot subtrees have the same structure but when using only matK, the conifer outgroup Podocarpus is nested deep within.

The similarity is indeed striking, in particular since the gene sample in the matrix comprises data from:
  • two of the nuclear-encoded ribosomal RNA genes (18S, 25S; biparentally inherited) that did follow partly different evolutionary trajectories, as e.g. well-studied in the case of Fagales (being a derived eudicot, not included in my matrix)
  • six chloroplast genes/gene regions (maternally inherited including the classics rbcL and matK but also the rpoC2, the most informative gene identified by Walker et al.)
  • three mitochondrial genes (also maternally inherited, but most mutations are, amino-acid-wise, synonymous, being concentrated at the third codon position).
The main things that matK get's wrong* in contrast to the combined tree are deep divergences represented by (very) short branches, in the part of the graph following the (very rapid) split of the mesangiosperm common ancestor (known as "Darwin's abominable mystery").

Also, it nests Podocarpus, the conifer in the outgroup, with unambiguous support in the monocots — which clearly is wrong, a false positive. Looking into the alignment, we can see that the reason for this is a mix of moderate-LBA (long-branch attraction) with missing-data-culling. To minimize LBA artifacts in the matrix originally used, I blanked out parts of the matK in the outgroup (which included a more derived conifer, Pinus, but also the extremely divergent gnetophytes); parts that were not straightforwardly alignable with the angiosperm matK.

The best way to illustrate internal signal conflict is, however, to directly show the BS Consensus network, not mapping support on two alternative topologies as seen in the tanglegram.

BS Consensus network based on 150 matK BS pseudoreplicates (numbers of necessary BS replicates determined by Pattengale et al.'s extended majority rule bootstop criterion implemented in RAxML)

When looking at BS << 100 and the boxes of competing splits in BS-support networks, it is important to keep in mind that low support can have two reasons:
  • Lack of decisive signal, because the BS pseudoreplicates will have (semi-)random or biased branching patterns; in the tree this surfaces usually as low (when random) to moderately high (when biased) support associated with (very) short branches.
  • Conflicting signals, ie. signals incompatible with a single tree; depending which site is eliminated or duplicated during resampling, the BS pseudoreplicate will show one or another topology; strong, deep conflict can surface in a tree by low support associated with (normally) long internal branches but also relatively high support for one alternative topology, the other only manifesting in very long terminal branches.
Regarding Walker et al.'s results, we now need to ask:
  1. Are the non-conflicted branches in the combined tree (major clades equal to the gold tree) the result of shared history of all of the included genes, or just that of the matK?
  2. Is the conflict with the combined tree and locally ambiguous signals due to a different history of the matK, located in the large single-copy unit, and the other genes, or just matK's inability to get certain things right?
In this case, all relatively high-supported conflicting matK splits are associated either with: (i) very short internal branches in the tree, the non-discriminative product of a fast ancient radiation, or (ii) are the result of an obvious data/branching artifact, ie. the misplaced Podocarpus.

So far, nothing challenges the assumption that the combined genes didn't follow the same history. Whether the other genes reveal something else, we'll see in my next post.



* or right: APG IV treats Ceratophyllum as the "probable sister of the eudicots" (see also Stevens' Angiosperm Phylogeny Website).

Monday, April 1, 2019

The Tree of Life (April 1)


The so-called Tree of Life is actually an anastomosing plexus rather than a divaricating tree, due to extensive interconnections between the cell and genome lineages during early single-cell evolution. These connections may have been caused by the process known as horizontal gene transfer.

Furthermore, the alleged Last Universal Common Ancestor may not have been a single coherent group, but may have been a mixture of quite different genotypes. After all, this supposed ancestor does not represent the origin of life, but was itself the end-product of an extensive prior evolutionary history.

These two basic points are illustrated in the following figure.


Happy April 1. For previous posts, see:

Monday, September 17, 2018

Getting the wrong tree when reticulations are ignored


One issue that has long intrigued me is what happens when someone constructs a phylogenetic tree under circumstances where there are reticulate evolutionary events in the actual (ie. true) phylogeny itself. That is, a network is required to accurately represent the phylogeny, but a tree is used as the model, instead. How accurate is the tree?

By this, I mean that, if the phylogeny can be thought of as a "tree with reticulations", do we simply get that tree but miss the reticulations, or do we get a different (ie. wrong) tree?


Sometimes, people refer to this situation as having a "backbone tree" — the phylogeny is basically tree-like, but there are a few extra branches, perhaps representing occasional hybridizations or horizontal gene transfers. The phylogenetic tree can then be treated as a close approximation to the true phylogeny, representing the diversification events but not the (rarer) reticulation events.

I have argued against this approach (2014. Systematic Biology 63: 628-638.). Instead of seeing a network as a generalization of a tree, we should see a tree as a simplification of a network. If we do this, then we would construct a network every time; and sometimes that network would be a tree, because there are no reticulation events in the phylogeny. It cannot work the other way around, because we can never get a network if all we ask for is a tree!

Presumably, if there are no reticulations then we should get the same answer (phylogenetic tree) irrespective of whether we simply construct a tree or instead construct a network that turns out to be a tree. But what about the "backbone tree" situation? Here, it has always seemed to me to be possible that we do not get the same tree. If this is so, then constructing a tree and then adding a few reticulations to it (as is often done in the literature) would not work — we would be adding reticulations to the wrong backbone tree.

There are two possible ways in which we can get the wrong backbone tree: the topology might be incorrect, or the branch-lengths might be incorrect (or both). For example, if there are true reticulations and yet we do not include them in our model, I have argued that the branches will be too short (2014. Systematic Biology 63: 847-849.) — two taxa will be genetically similar because of the reticulation events, but the tree-building algorithm can only make them similar on the tree by shortening the branches (not by adding a reticulation).

Fortunately, for at least one tree-building model Luay Nakhleh and his group have now done some simulations to answer my questions. You may not yet have noticed their results, because they are not necessarily in the most obvious place; so I will highlight them here. The analyses involve the Multispecies Coalescent (MSC) model, which accounts for incomplete lineage sorting during the tree-like part of evolution, as compared to the Multispecies Network Coalescent (MSNC) which adds reticulations (eg hybridization) to the model.

1.
Dingqiao Wen, Yun Yu, Matthew W. Hahn, Luay Nakhleh (2016) Reticulate evolutionary history and extensive introgression in mosquito species revealed by phylogenetic network analysis. Molecular Ecology 25: 2361-2372.

This paper compares a tree-based analysis (construct a tree first then add reticulations) with a network-based analysis (construct a network) for an empirical genomic dataset. The two results differ.

2.
Dingqiao Wen, Luay Nakhleh (2018) Coestimating reticulate phylogenies and gene trees from multilocus sequence data. Systematic Biology 67: 439-457.

Tucked away in the Supplementary Information are the results of a set of simulations comparing the MSC (using *Beast) and the MSNC (using PhyloNet), with (section 3) and without (section 2) reticulations. The basic conclusion is that, in the presence of reticulation, tree-based methods either get the tree completely wrong, or they get the tree topology right but the branch lengths are "forced" to be very short. A summary of the latter result is shown in the figure above. In the absence of reticulation, both methods produce the same tree.

3.
R.A. Leo Elworth, Huw A. Ogilvie, Jiafan Zhu, and Luay Nakhleh (ms.) Advances in computational methods for phylogenetic networks in the presence of hybridization. (chapter for a forthcoming book]

A summary of the group's work to date. Section 6.3 summarizes the results from the paper 2.

Tuesday, July 18, 2017

Stacking neighbour-nets: ancestors and descendants


Although they are not phylogenetic networks in an evolutionary sense, neighbour-nets are amazingly efficient when it comes to depicting actual ancestor-descendant relationships. Spencer et al. (2004), for example, applied various methods of phylogenetic inference to reconstruct a known phylogeny of scripts, text copies made by scribes based on an original text, or copies of that text. They found that the neighbour-net algorithm produced a graph that best depicts the actual ancestor (original text) to descendant (text copy made by a scribe) relationships. The reason is relatively simple: neighbour-nets are well-suited to extract and reflect differentiation patterns from a distance matrix.

In this post I will explore this idea in some detail, because I think that it has important practical implications (which I will further elaborate in future posts). If data are available for time slices of the evolutionary history, then it is possible to stack a series of networks that describe each time slice, thus providing a much more comprehensively inferred genealogy.

Neighbor-nets

If a distance matrix reflects exactly the phylogeny (i.e. if the signal from the matrix is trivial), then the neighbour-net looks like a tree, with the ancestors placed seemingly at the internal nodes (medians) of the subtree(s) containing their descendants — that is, the neighbour-net looks like a median network (Figure 1). In fact, this is just mimicry. The ancestors are not actually placed on the internal nodes (medians) but are connected by zero-length edge bundles to the centre of the graph (the roots of their descendants).


Figure 1. Neighbour-net inferred from a perfect distance matrix (from Denk & Grimm 2009)

In reality, distance matrices will not be exact reflections of the phylogeny (the ‘true’ tree) but distorted; and this will be reflected also in the neighbour-net (Fig. 2).


Figure 2. Neighbour-net inferred from an imperfect distance matrix (cf. Denk & Grimm 2009, fig. 1)

I superimposed a potentially inferred tree on Figure 2 to highlight some distorting effects. Note that this could be a tree optimised using a distance criterion such as minimum evolution or least-squares, or a tree optimised under maximum likelihood or maximum parsimony (one of the alternative topologies found in the sample of equally parsimonious trees).

The following distorting effects may apply during tree-inference: (i) a misplaced outgroup-inferred ingroup root (due to convergences shared by the outgroup and members of lineage B), (ii) lineage B is dissolved into a grade, although it should be a clade (convergences shared by members of lineages A and B), and (iii) the all-ingroup ancestor is resolved as the sister to lineage A, but not B. The neighbour-net illustrates the uncertainties of placing several taxa (the box-like parts of the graph), while keeping the ancestors equally close to their descendants. This provides information lost (or overlooked) when just inferring a tree (but not by exploratory data analysis of e.g. the bootstrap support patterns).

The basics — why stack networks?

Figure 3 shows a hypothetical evolution of a phylogenetic lineage in a two-dimensional morpho-space. The common ancestor (black dot) gives rise to two, morphologically somewhat distinct lineages (bluish vs. reddish coloured). The lineages evolve over time and diverge again. The overall differentiation within the group increases, but eventually the potential niche/morphospace is filled. When looking only at the final situation (i.e. the modern-day situation), we may be tempted to infer wrong relationships based on the morphological distinctness. Each one of the blue and red daughter lineages evolved into similar niches and obtained somewhat similar morphological character suites substantially distinct from the one of their respective sister taxa. Translating this situation into a distance matrix (or a character matrix) to infer a tree, will provide a wrong topology, when long-branch attraction steps in, that recognises one or two blue-red sister pair(s).


Figure 3. Evolution (vertically) and diversification of a lineage in a two-dimensional morphospace (horizontally)

Adding all of the ancestors can help to escape long-branch attraction (Figure 4; see also Wiens 2005). We don’t find red111 as sister to blu11, but still resolve the wrong sister relationship between the overall too-similar endpoints of the converging red and blue sub-lineages (red222 and blu22). The remainder of the red lineage is erroneously dissolved into (A) an ancient sister lineage (red0); (B) a real clade (red1, red11, red111), the red sub-lineage most distinct from the blu lineage; and (C) a grade (red2, red22), “basal” (wrong terminology, but often used) to the blue lineage, collecting the older members of the red sub-lineage evolving towards the morphospace of the blue lineage.


Figure 4. Neighbour-joining tree inferred on a distance matrix exactly reflecting the pairwise distances along both axes in the 2-dimensional morphospace

Only a few, (data-wise) trivial branches of the true tree (broadened green edges) are found in the inferred tree. An obvious defect of this tree is that the phylogenetic distance, the sum of branch lengths between two tips, does not reflect the pairwise distances encoded in the matrix. For instance, the all-ancestor should be equally distant from both of the ancestors of the red and blue lineages (red0, blu0).

The neighbour-net shows something rather different (Figure 5). A large box can be seen, referring to the highly incompatible signal induced by the ancestors and their direct and subsequent descendants.


Figure 5. Neighbour-net splits graph inferred on the same distance matrix

Red222 and blu22, the false sisters, are placed next to each other and share an edge bundle. They are most similar to each other and increasingly distinct from all other taxa included in the analysis. However, their nearest relatives are their actual ancestors (red22 and blu2), which are already quite distinct from each other, and show affinities to further members to their clade but not to the other clade (the topology of the tree from Figure 4 is shown in yellow in Figure 5).

The network includes the true tree in addition to edge bundles referring to wrong alternatives (induced by imperfect data; in this case, convergence due to evolution into a similar morphospace). The phylogenetic distances between two tips (via alternative pathways) reflect much better the pairwise distances (almost exactly). Thus, the neighbour-net is a much more comprehensive display of the actual signals in the imperfect matrix than a tree could ever be (independent of the optimisation criterion used).

The fossils included in our example represent a time sequence, an actual change in time. With that information as background, we can much more easily access the neighbour-net’s structure (Figure 6).


Figure 6. The same neighbour-net with time-slices

Already in the second time-slice the lineage diverged into two distinct branches represented by red0 and blu0. Both evolved (blu0→blu00) and diverged (red1, red2) in the next time slice, and so on. The fact that the neighbour-net is not a phylogenetic graph (in an evolutionary sense) becomes a strength. In any tree, we would need to deduce, or be tempted to deduce, (inclusive) common origins (Hennig’s monophyly) from the clades exhibited in the rooted version of the tree. Here, using the most natural root to root the tree, the oldest representative of the lineage (ancestor), misplaces the two representatives of the second time-slice along with those involved in subsequent radiations.

Stacking networks

There are two simple ways to trace the change in differentiation patterns through time using stacked networks: (A) Generate a series of networks per time-slice and identify the closest relatives in the next (and/or preceding) time-slice, or (B) Generate networks that combine the taxa of two subsequent time-slices.


Figure 7. A sequence of neighbour-nets, with the taxa filtered by age. (These are actual reconstructions inferred by SplitsTree based on the taxon-filtered subsets of the all-inclusive distance matrix.)

My example only includes up to four co-eval taxa, and hence the neighbour-nets are trivial graphs for each time-slice (Figure 7). The connecting lines between the time-slices indicate the closest and next-closest possible descendants (as defined by the smallest and next-smallest distance) of each earlier taxon. The thickness of the connections reflects their absolute similarity — the thickest lines indicate a morphological pairwise distance of 0.13, and the thinnest are distances of > 0.33. The colour indicates whether the connection reflects a true (green) or false (yellow) relationship.

Analysed this way, the matrix’ signal appears quite perfect in relation to the true tree. One can trace the increasing diversity within the clade (all descendants of the all-ancestor), as well as the misleading decreasing distance between the blu2/blu22 and red22/red222 lineages. In this case, the same stacking procedure would also work with trees, as the neighbour-nets are trivial and very tree-like. With real world data, the differences may be more profound.

With more complex or less complete data, we have a higher risk that ancestor-descendant relationships will not be straightforwardly identified by highest-similarity pairs. Missing data, for instance, can result in distances that mask or over-estimate the actual phylogenetic distances between an older and a younger taxon. Using the stacking procedure illustrated in Figure 7, such problems can become visible in the form of related taxa from the same time-slice that are connected to unrelated or distantly related taxa in the preceding or following time-slices.

But how to identify more likely candidates? One possibility is to assess potential phylogenetic relationship by combining taxa of two subsequent time-slices. The connectives between the reconstructions are then straightforward: each taxon is always used in two different reconstructions (Figure 8). This procedure allows us to establish the phylogenetic affinities of a taxon with respect to co-eval and older taxa (potential siblings and ancestors) or co-eval and younger taxa (potential siblings and descendants).


Figure 8. A sequence of neighbour-nets, each one including the taxa of two subsequent time-slices


Stacking networks — what’s next?

The first thing, obviously, is to test the suggested procedures for real-world data, involving groups with a dense and well-studied fossil record. I will provide a real-world example in my next post using the matrix we put up for our systematic revision of Osmundaceae (King Ferns) rhizomes (Bomfleur, Grimm & McLoughlin 2017).

Simulations may help to identify misleads caused by missing data, and the resulting distorted distance matrices, and non-comprehensively sampled (time-wise) phylogenies. They may also be informative regarding whether consensus networks reflecting competing branch support can be used for similar approaches.

Programmers are needed, too. For my graphics, I established the inter-time-slice connectives by hand; but it would be handy to have a programme environment that can do this.

References

Bomfleur B, Grimm GW, McLoughlin S. 2017. The fossil Osmundales (Royal Ferns)—a phylogenetic network analysis, revised taxonomy, and evolutionary classification of anatomically preserved trunks and rhizomes. PeerJ 5: e3433. Open access: https://peerj.com/articles/3433/

Denk T, Grimm GW. 2009. The biogeographic history of beech trees. Review of Palaeobotany and Palynology 158: 83-100.

Spencer M, Davidson EA, Barbrook AC, Howe CJ. 2004. Phylogenetics of artificial manuscripts. Journal of Theoretical Biology 227: 503-511.

Wiens JJ [, Soltis P ?]. 2005. Can incomplete taxa rescue phylogenetic analyses from long-branch attraction? Systematic Biology 54: 731-742. Open access: https://academic.oup.com/sysbio/article-lookup/doi/10.1080/10635150500234583

Tuesday, May 16, 2017

Connecting tree and network edges


I have struggled over the years to try to understand the relationship between trees and networks. In one sense, networks are generalizations of trees, and in another sense a tree is just a simplified network. But it is not always that simple.

For example, not all networks can be created by adding edges to a tree (see Networks vs augmented trees); so the connection between trees and networks is not always obvious. Moreover, it is not always easy to determine which tree edges are present in any given network, or which network edges are present in a given tree.

Nevertheless, this should be basic information in phylogenetics — otherwise, how can we know when a tree is adequate for our purposes, or when a network is needed?

It turns out that I have not been alone in struggling to connect trees and networks. Fortunately, some of these other people decided to actually do something about it, rather than simply struggling on. As a result, a computerized way to relate much of the important information connecting trees with networks now exists.
Klaus Schliep, Alastair J. Potts, David A. Morrison and Guido W. Grimm
Intertwining phylogenetic trees and networks.
Methods in Ecology and Evolution (Early View)
To quote the authors:
Here we provide a framework, implemented in the PHANGORN library in R, to transfer information between trees and networks. This includes: (i) identifying and labelling equivalent tree branches and network edges, (ii) transferring tree branch-support to network edges, and (iii) mapping bipartition support from a sample of trees (e.g. from bootstrapping or Bayesian inference) onto network edges.
These three functions are illustrated in this figure, taken from the paper. It should be self-explanatory to anyone who has tried to relate the edges of trees and networks; but if it is not, then you can read an explanation in the paper.


The R library referred to, including the source code, along with some examples and vignettes, can be accessed on the PHANGORN CRAN page.

Note that PHANGORN (originally created by Klaus Schliep) also contains other functions related to estimating phylogenetic trees and networks, using maximum likelihood, maximum parsimony, distance methods and hadamard conjugation. Specifically, it allows you to: estimate phylogenies, compare trees and models, and explore tree space and visualize phylogenetic trees and split graphs.

Tuesday, April 18, 2017

Multimedia phylogeny?


Evolutionary concepts have often been transferred to other fields of study, or derived independently in them, especially in anthropology in the broadest sense, covering all cultural products of the human mind. This includes phylogenetic studies of languages, texts, tales, artifacts, and so on — you will find many examples of such studies in this blog. One of the more recent applications has been to what is sometimes called multimedia phylogeny — the research field that "studies the problem of discovering phylogenetic dependencies in digital media".

I have noted before that phylogenetics in the biological sense is an analogy when applied to other fields, because only in biology is genetic information physically transferred between generations — in the other fields, cultural information transfer is all in the minds of the people, not in their genes (see False analogies between anthropology and biology). This analogy often becomes problematic when applied to other fields, because the practical application of bioinformatics techniques separates the informatics from the bio, and the mathematical analyses focus on trying to implement the informatics without any biological justification.


A recent paper that discusses the application of bioinformatics to multimedia phylogeny exemplifies the potential problems:
Guilherme D Marmerola, Marina A Oikawa, Zanoni Dias, Siome Goldenstein, Anderson Rocha (2017) On the reconstruction of text phylogeny trees: evaluation and analysis of textual relationships. PLoS One 11(12): e0167822.
The authors described their background information thus:
Articles on news portals and collaborative platforms (such as Wikipedia), source code, posts on social networks, and even scientific publications or literary works, are some examples in which textual content can be subject to changes in an evolutionary process. In this scenario, given a set of near-duplicate documents, it is worthwhile to find which one is the original and the history of changes that created the whole set. Such functionality would have immediate applications on news tracking services, detection of plagiarism, textual criticism, and copyright enforcement, for instance.
However, this is not an easy task, as textual features pointing to the documents' evolutionary direction may not be evident and are often dataset dependent. Moreover, side information, such as time stamps, are neither always available nor reliable. In this paper, we propose a framework for reliably reconstructing text phylogeny trees, and seamlessly exploring new approaches on a wide range of scenarios of text reusage. We employ and evaluate distinct combinations of dissimilarity measures and reconstruction strategies within the proposed framework.
So, their solution to the separation of bio from informatics is to try a range of techniques, none of which are based on any particular model of how phylogenetic changes might occur in text documents. All of these methods involve distance-based tree-building.

The essential problem, as I see it, is that without a model of change there is no reliable way to separate phylogenetic information from any other type of information. For example, similarity can arise from many sources, only some of which provide information about phylogenetic history — phylogenetic similarity is a form of "special similarity". In biology, other sources of similarity are usually lumped together as chance similarities, such as convergence, parallelism, etc. Without this basic separation of phylogenetic and chance similarity, it does not matter how many distance measures you use, or how many tree-building methods you employ — if you can't separate phylogeny from chance then you are wasting your time constructing a hypothetical  evolutionary history.

The authors' only saving grace is their claim that: "In text phylogeny, unlike stemmatology [the analysis of hand-written rather than digital texts], the fundamental aim is to find the relationships among near-duplicate text documents through the analysis of their transformations over time." The expectation, then, is that the phylogenetic similarity of the texts will be high, which will thus reduce the possibility of chance similarities. Sadly, it will also reduce the probability that the similarities will contain any phylogenetic information at all — this is the classic short-branches-are-hard-to-reconstruct problem in phylogenetics.

For digital texts, the authors employ three distance measures: edit distance, normalized compression distance, and cosine similarity. None of these are model-based in any phylogenetic sense (although the first one is used in alignment programs such as Clustal) — I have discussed this in the post on Non-model distances in phylogenetics. Their tree-building methods include: parsimony, support vector machines (a machine-learning form of classification), and random forests (a decision-tree form of classification). Once again, none of these is model-based in terms of textual changes.

A final issue is the insistence on trees as the model of a phylogeny. In stemmatology, for example, a network is a more obvious phylogenetic model, because hand-written texts can be copied from multiple sources. Indeed, this distinction plays an important role in the first application of phylogenetics to stemmatology (see the post on An outline history of phylogenetic trees and networks). Perhaps this is not an issue for "near-duplicate text documents", but it does seem like an unnecessary restriction. Moreover, one of the empirical examples used in the paper actually has a network history, which therefore does not match the authors' reconstructed tree.

Tuesday, January 10, 2017

Why do we need Bayesian phylogenetic information content?


There are many ways to construct a phylogenetic tree, and after we have done so we are usually expected to indicate something about "branch support", such as bootstrap values or bayesian posterior probabilities. Rarely, however, do people indicate whether there is much tree-like phylogenetic information in their dataset in the first place — it is simply assumed that there must be (fingers crossed, touch wood).

Recently, this latter issue has been addressed for bayesian analysis by:
Paul O. Lewis, Ming-Hui Chen, Lynn Kuo, Louise A. Lewis, Karolina Fučíková, Suman Neupane, Yu-Bo Wang, Daoyuan Shi. (2016) Estimating Bayesian phylogenetic information content. Systematic Biology 65: 1009-1023.
They develop a methodology for "measuring information about tree topology using marginal posterior distributions of tree topologies", and apply it to two small empirical datasets. That is, we can now work out something about "[substitution] saturation and detecting conflict among data partitions that can negatively affect analyses of concatenated data."

However, we have long been able to do this with data-display phylogenetic networks. More to the point, we can do it in a second or two, without ever constructing a tree. More pedantically, if the network construction produces a tree, then we know there is tree-like phylogenetic information in the dataset; if we get a network then there is little such information. Equally importantly, the network might tell us something about the patterns of non-tree-likeness, which a single-number measurement cannot.

Let's take the first empirical dataset, as described by the authors:
The five sequences of rpsll composing the data set BLOODROOT [three taxa from the angiosperm family Papaveraceae and two monocots] ... were chosen because they represent a case in which horizontal transfer of half of the gene results in different true tree topologies for the 5′ (219 nucleotide sites) and 3′ (237 nucleotide sites) subsets, which allows investigation of information content estimation in the presence of true conflicting phylogenetic signal. We analyzed each half of the data separately and measured phylogenetic dissonance, which is expected to be high in this case.
Here is the NeighborNet based on uncorrected distances. The idea that there is something non-tree-like about Sanguinaria seems hard to avoid. Indeed, the network pattern makes recombination an obvious first choice, with part of the sequence matching the Papaveraceae (on the left) and part matching the monocots (on the right). This recombination may be due to HGT.


Now for the second dataset:
The data set ALGAE comprises chloroplast psaB sequences from 33 taxa of green algae (phylum Chlorophyta, class Chlorophyceae, order Sphaeropleales) ... The alignments of just the psaB gene ... were chosen because of their deep divergence, which invites hasty judgements of saturation, especially of third codon position sites. We analyzed second and third codon position sites separately ... to assess which subset has more phylogenetic information.
Here are the two NeighborNets based on uncorrected distances. Once again, it is immediately obvious that the third-codon positions have almost no information at all, even for a network, let alone a tree — the terminal branches do not connect in any coherent way. The second-codon positions do have some information, but it is so contradictory that one could not construct a reliable tree. Saturation of nucleotide substitutions is a likely candidate for this situation; and some correction for this saturation would be needed even to construct a reasonable network from these data.

2nd positions:

3rd positions:

Tuesday, December 6, 2016

Why are splits graphs still called phylogenetic networks?


This is an issue that has long concerned me, and which I think causes a lot of confusion among biologists. A phylogenetic tree is usually a clear concept — to a biologist, it is a diagram that displays a hypothesis of evolutionary history. The expectation, then, is that a phylogenetic network does the same thing for reticulate evolutionary histories. However, this is not true of splits graphs; and so there is potential confusion.

Mathematically, of course, a phylogenetic tree is a directed acyclic line graph. It is usually constructed, in practice, by first producing an undirected graph based on some pattern-analysis procedure, and then nominating one of the nodes or edges as the root (say, by specifying an outgroup). So, the mathematics is not really connected to the biological interpretation. To a mathematician, the tree is a set of nodes connected by directed edges, and the nodes could represent anything at all, as could the edges. It is the biologist who artificially imposes the idea that the nodes represent real historical organisms connected by the flow of evolution — ancestors connected to descendants by evolutionary events.

A phylogenetic network should logically be a generalization of this idea of a phylogenetic tree, adding the possibility of evolutionary relationships due to gene flow, in addition to the ancestor-descendant relationships. This can be done, but it is only partly done by splits graphs.

That is, a splits graph generalizes the idea of an undirected line graph (an unrooted tree), but not a directed acyclic graph (a rooted tree). It follows the same logic of using a pattern-analysis procedure to produce an undirected graph, although the graph can have reticulations, and thus is a network rather than necessarily being a bifurcating tree. However, it is not straightforward to specify a root in a way that will turn this into an acyclic graph. So, in general it does not represent a phylogeny.

Indeed, splits graphs are simply one form of multivariate pattern analysis, along with clustering and ordination techniques, which are familiar as data-display methods in phenetics (see Morrison D.A. 2014. Phylogenetic networks — a new form of multivariate data summary for data mining and exploratory data analysis. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 4: 296-312). In this sense, it makes no difference whatsoever what the data represent — they can be data used for phylogenetics, or they could be any other form of multivariate data. Indeed, this point is illustrated in many of the posts in this blog, which can be accessed in the Analyses page.

So, unlike unrooted trees, unrooted splits graphs are not a route to producing a phylogenetic diagram. Mind you, they are a very useful form of multivariate data analysis in their own right, and I value them highly as a form of exploratory data analysis. But that doesn't make them phylogenetic networks in the biological sense.

So, isn't it about time we stopped calling splits graphs "phylogenetic networks"? They aren't, to a biologist, so why call them that?

Tuesday, June 21, 2016

Alignments and phylogenetic reconstruction in linguistics and biology


In a very interesting article from 2009 (Morrison 2009), David discusses the question of why phylogeneticists would "ignore computerized sequence alignment". This article was really interesting to me for two reasons: First, the article provides some interesting statistics regarding the degree to which biologists manually adjust the alignments that were automatically produced by software. Second, the article points to the seemingly strange situation in biology in which tree-building is considered to be a task that can be entirely carried out by machines, while the majority of scholars would not trust their final sequence alignments to a computer (Morrison 2009: 150).

This situation finds a direct analogon in historical linguistics. Phylogenetic reconstruction is gaining more and more ground, with many scholars applying (mostly Bayesian) phylogenetic tools to analyze their data (Indo-European: Bouckaert et al. 2012, Tupí-Guaraní (South America): Michael et al. 2015, Japonic: Lee and Hasegawa 2011, Pama-Nguyan (Australian): Bowern and Atkinson 2012, Semitic: Kitchen et al. 2009, Bantu: Grollemund et al. 2015, etc.). Fully automated workflows involving automatic sequence comparison are also practiced (Holman et al. 2011, Jäger 2015, Wheeler 2015), but many linguists remain sceptical regarding their results.

One major difference between biology and linguistics is the selection of comparanda. Biological methods usually derive phylogenetic trees from multiply aligned sequences. Linguistic methods derive trees from sets of homologous (cognate) words (cognate sets) distributed across languages whose evolution is modeled as a process of word-gain and word loss (similar to gene-family gain-loss-studies in biology). While biologists fiddle with their alignments, linguists fiddle with their cognate sets. Cognate identification is exclusively done manually at the moment, and scholars use all kinds of information about word relations that they can get, be it etymological dictionaries, which have been published for more than 200 years, or the intuition of the expert who is annotating the data for cognacy.

Identification of cognate sets in linguistics is essentially a task of sequence comparison (List 2014), and algorithmic as well as manual procedures involve the multiple and the pairwise alignment of words (even if it is done only implicitly by human experts). Compared to biology, sequence comparison in historical linguistics is exacerbated by two factors:
  • alphabets (phoneme systems) in linguistics are themselves mutable (Geisler and List 2013), so that when aligning two words we need to find both a mapping between the two alphabets, translating one alphabet into the other, plus a scoring function by which we can score the alignment,
  • regular sound change (the process by which the phoneme system is changed) and sporadic sound change (the process by which a sound is sporadically assimilated, lost, or added) are not the only processes that contribute to change of words in the lexicon, and morphological change (by which whole blocks of meaningful parts of a word are re-arranged, exchanged, lost, or added) yields patterns that are essentially unalignable.
The problem of finding the correct mapping between two alphabets in linguistics is further exacerbated by language contact: If languages exchange words on a large scale, then this may have a huge impact on the system of the languages, and it may even introduce new sounds to a language that were not there before (thanks to English, German has now the sound [dʒ], as in journalist or job). If borrowing is frequent enough, it may get close to impossible to judge from comparing the words alone, whether two words in different languages have been transferred directly (vertically) from an ancestral language, or laterally.

As a result, it is probably understandable why linguists often refuse to carry out full alignments of the words in their data. An alignment itself does not necessarily tell us much, compared to all of those processes that an expert infers when comparing language data, which are not alignable.

As an example, let us consider the word for "sun" in six Indo-European languages. Since "sun" is a very basic concept, probably fundamental for all human cultures, experts assume that this word was present as *séh₂u̯el- in Indo-European (an asterisk indicates that the word is not reflected in written sources), and that it was retained as Russian солнце [sɔnʦə], Polish słońce [swɔnjʦɛ], French soleil [sɔlɛj], Italian sole [sole], German Sonne [sɔnə], and Swedish sol [suːl] (Wodtko et al. 2008). An obvious alignment, reflecting the surface similarity between all of these words, would be the following one (taken from List 2014: 135):

Alignment based on sequence similarity.

This alignment, however, is by no means correct. Russian [sɔnʦə] and Polish [swɔnʲʦɛ], for example, share a common suffix, which is reflected as [nʦə] in Russian and as [nʲʦɛ] in Polish, and which was innovated in the the common ancestor of Russian and Polish, but is not present in either of the four other languages. So the [n] in German [sɔnə] is essentially not homologous with the [n] in Russian or the [nʲ] in Polish. The same applies to the [ɛj] in French [sɔlɛj] which reflects a diminutive suffix in Latin sol-iculus "small sun", the regular ancestor form of French soleil. Furthermore, the [w] in the Polish word regularly corresponds to the [l] in French, Italian, and Swedish, but it reflects a swap (metathesis) in the order of the vowel and the consonant in Polish — [sɔl] became [slɔ] which became [swɔ]).

Taking all (and more) of this into account, we need to modify our alignment to account more closely for the processes that experts have inferred from intensive language comparison, as shown in the next figure below (taken from List 2014: 135). In this alignment, the swap in Polish is reflected by the white font of the sounds involved, and gray-shaded columns are supposed to reflect the oldest layer of homology.

Historically informed alignment.

However, even this alignment is essentially misleading. The Indo-European word for "sun" supposedly had a complex paradigm in which the word's stem was alternating in the nominative (and accusative) case and the other cases (oblique cases). So, nominative and accusative used the stem *sóh₂u̯el-, while the other cases used the stem *sh₂én-. The Russian, Polish, French, Italian, and the Swedish form go back to the former, while the German form goes back to the latter, since it is further assumed (or it can be assumed) that the alternation was still preserved in the ancestor of Swedish and German.

This means, however, that our alignment above shrinks to an alignment in which only the first letter, the s, is still reflected in all languages! The following graphic (taken from List 2016) illustrates the processes that led to the current situation for four of our six languages:

Morphological processes of lexical change.

What does this example tell us? On the one hand, it gives some explanation for why linguists do not really want to align words (although the first alignments go back to the early 20th centur, cf. Dixon and Kroeber 1919). It also explains, why classical linguists have a very sceptical attitude towards the computerization of word comparisons, based on the (partially justified) assumption that computers could not handle the complex patterns that are so characteristic of language change.

On the other hand, comparing the situation with biology as reported in Morrison (2009), we can find an interesting parallel between the two disciplines: both linguists and biologists do not really trust machines for comparing their sequences (albeit at different levels of analysis), but they do not seem to have many problems in trusting machines to reconstruct their trees.

However, especially this last point, the fact that we trust machines to grow our trees, while we distrust them to prepare the seeds, should ring an alarm bell. First, we seem to lack clear guidelines (at least in linguistics) regarding the way the manual adjustment (of alignments in biology and cognate sets in linguistics) should be carried out, which has a clear impact on repeatability. Second, if we have processes in both fields that yield essentially unalignable patterns, such as duplications and other molecular processes in biology (Morrison 2009: 156), and morphological processes in linguistics, how can we assume that a phylogenetic tree analysis can sufficiently cope with them, even if we manually adjust everything?

References
  • Bouckaert, R., P. Lemey, M. Dunn, S. Greenhill, A. Alekseyenko, A. Drummond, R. Gray, M. Suchard, and Q. Atkinson (2012): Mapping the origins and expansion of the Indo-European language family. Science 337.6097. 957-960.
  • Bowern, C. and Q. Atkinson (2012): Computational phylogenetics of the internal structure of Pama-Nguyan. Language 88. 817-845.
  • Dixon, R. and A. Kroeber (1919): Linguistic families of California. University of California Press: Berkeley.
  • Geisler, H. and J.-M. List (2013): Do languages grow on trees? The tree metaphor in the history of linguistics. In: Fangerau, H., H. Geisler, T. Halling, and W. Martin (eds.): Classification and evolution in biology, linguistics and the history of science. Concepts – methods – visualization. Franz Steiner Verlag: Stuttgart. 111-124.
  • Grollemund, R., S. Branford, K. Bostoen, A. Meade, C. Venditti, and M. Pagel (2015): Bantu expansion shows that habitat alters the route and pace of human dispersals. Proceedings of the National Academy of Sciences 112.43. 13296–13301.
  • Holman, E., C. Brown, S. Wichmann, A. Müller, V. Velupillai, H. Hammarström, S. Sauppe, H. Jung, D. Bakker, P. Brown, O. Belyaev, M. Urban, R. Mailhammer, J.-M. List, and D. Egorov (2011): Automated dating of the world’s language families based on lexical similarity. Curr. Anthropol. 52.6. 841-875.
  • Jäger, G. (2015): Support for linguistic macrofamilies from weighted alignment. Proceedings of the National Academy of Sciences 112.41. 12752–12757.
  • Kitchen, A., C. Ehret, S. Assefa, and C. Mulligan (2009): Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East. Proc. R. Soc. London, Ser. B 276.1668. 2703-2710.
  • Lee, S. and T. Hasegawa (2011): Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages. Proc. R. Soc. London, Ser. B 278.1725. 3662-3669.
  • List, J.-M. (2014): Sequence comparison in historical linguistics. Düsseldorf University Press: Düsseldorf.
  • List, J.-M. (2016): Beyond cognacy: Historical relations between words and their implication for phylogenetic reconstruction. Journal of Language Evolution 1. DOI: 10.1093/jole/lzw006.
  • Michael, L., N. Chousou-Polydouri, K. Bartolomei, E. Donnelly, V. Wauters, S. Meira, and Z. O’Hagan (2015): A Bayesian phylogenetic classification of Tupí-Guaraní. LIAMES 15.2. 193-221.
  • Morrison, D. (2009): Why would phylogeneticists ignore computerized sequence alignment? Syst. Biol. 58.1. 150-158.
  • Wheeler, W. and P. Whiteley (2015): Historical linguistics as a sequence optimization problem: the evolution and biogeography of Uto-Aztecan languages. Cladistics 31.2. 113-125.
  • Wodtko, D., B. Irslinger, and C. Schneider (2008): Nomina im Indogermanischen Lexikon [Nouns in the Indo-European lexicon]. Winter: Heidelberg.