Showing posts with label Cladistics. Show all posts
Showing posts with label Cladistics. Show all posts

Monday, January 6, 2020

Why we may want to map trait evolution on networks, pt. 1 – Introduction


One of the more interesting aspects of studying evolution is to trace the evolution of the traits possessed by the organisms, whether those traits are physical or not (such as languages). That is, we usually infer phylogenetic trees and networks to see how things evolve, including both the organisms and their characteristics. However, this can easily lead to circular reasoning, as I will discuss here.

Background

A phylogenetic tree may be enough to work out who is sister to whom. However, when thinking about evolution itself, we actually want to find out who comes from whom, instead. This may be the reason why Charles Darwin did not title his 'abstract' A Natural Order of Species but used instead The Origin of Species.

The tricky bit is this: in order to find the origin, we first need to establish ancestor-descendant relationships, so that we can then see how things like fossils fit in (ie. whether they represent extinct sister lineages or precursors of the modern-day taxa). When taxon B is a derivation of A (ie. B evolved from A), the character suite of A is not only primitive but also the original set. Now, let's assume that we have a third fossil taxon C, which is clearly related to A and B. As evolutionary biologists, we cannot be content merely with inferring sister relationships between A, B, and C, but instead we need to decide whether or not C is also descended from A.

Ironically (from a modern cladistic viewpoint, focusing on establishing sister-clade relationships), Willy Hennig provided us with some tools for doing this:
  • We all know that apomorphies are derived traits, either unique (aut-) or uniquely shared by a group with inclusive common origin (syn-). Aut- and syn-apomorphies define (Hennig's) monophyletic groups. According to Hennig, a synapomorphy is a necessary criterion for recognizing a monophyletic group, and also a sufficient one (although the latter easily falls prey to circular reasoning). More importantly, they tell us that the ancestor(s) of the group and its (potentially lost) sister lineages lacked this trait!
  • Sym-plesiomorphies are traits that are primitive within a certain lineage. They define paraphyletic groups, which are groups of exclusive common origin. Following Hennig, they need to be discarded for systematics. From an evolutionary viewpoint, however, symplesiomorphies have double information content: (a) they provide us with traits that, at some point back in time, were synapomorphies; and (b) any member of a lineage not carrying the symplesiomorphic trait, shows a derived one.
Farris' cladistics is still the basis for systematics, and is widely applied in phylo-paleontology. The initial flaw of this approach is to assume that we can use morphological traits to infer a tree (with parsimony), and then map the same traits onto the inferred tree, allowing us to qualify the traits towards Hennig's objectives. How can this not be circular reasoning? We are mapping the traits onto a tree derived from those traits in the first place, so that the tree-building and mapping are not independent.

A simple (?) real-world example

For the purpose of this exercise, we will take the minute (seven character) matrix of Wilf et al. (2019) from this previous blog post (characters 5 and 7 corrected, and missing Fagaceae added; see also Denk et al. 2019, Science, 10.1126/science.aaz2189).

Wilf et al. found an Eocene fossil in South America, and argued that it must be a member of the modern genus Castanopsis, based on a parsimony DNA-scaffold approach (without actually using a DNA partition). Being a member of a modern genus, the fossil should have some aut-/synapormphies or at least symplesiomorphies or homoiologies characterizing its sublineage of the Fagaceae, the paraphyletic Castanoideae.

Based on the morphology, we can infer this tree:

Fig. 1 – Adams consensus tree of 3 most-parsimonious trees (11 steps, CI = 0.84, RI = 0.88), traits are mapped using Mesquite's default parsimony model. Castanopsis rothwellii is the Eocene fossil found by Wilf et al.

Two characters qualify as near-synapomorphies (effectively there is only one: hemispheric indehiscent cupules) that define a crown-clade including Lithocarpus, Notholithocarpus, Castanopsis (as part of intrageneric variation) and Quercus. Most other putatively derived traits within the Fagaceae subclades are symplesiomorphies; two are potential homoiologies, one defining the Castanea-Chrysolepis clade. [Note the staircase-like tree topology, a common feature of parsimony trees dealing with extinct lineages.] The fossil's character suite is relatively derived, characters 6 (shared only with some Castanopsis) and 7 (reversal as in Quercus) could be interpreted as an extinct side lineage of the (paraphyletic) Castanoideae.

This is not a bad analysis for seven characters, but it is likely to be quite wrong.

Fagaceae still exist today, and their DNA can be sampled. Below is a maximum-likelihood tree, based on a 2012 NCBI GenBank oligogene data harvest I did for a talk in Bordeaux — the alignment is 19,242 basepairs long, has 2,985 distinct alignment patterns and a gappyness of 35.8%. Each genus and major intra-generic lineage is represented by a strict consensus sequence based on all available data (checked for mislabeled or pseudogene accessions). [Oaks started to radiate > 50 Ma, Grímsson et al. 2015, Hipp et al. 2019; beeches about the same time, Denk et al. 2009, Renner et al. 2016.]

Fig. 2 – a ML tree based on strict genus/intrageneric consensus sequences (see also Oh & Manos, 2008, fig. 4, based only on data from the Crabs Claw gene, CRC; fig. 5 in the same paper shows a combined CRC + ITS tree)

According to this analysis, Chrysolepis and Castanea are not sisters; Castanea, but not Lithocarpus, is a close relative of the oaks. The (monophyletic) Trigonobalanoideae should form a clade (Fig. 2) not a grade (Fig. 1).

The analysis is not circular anymore, when we infer a tree based on data that is, as far as we know, independent of the data we want to map onto the tree. With the invention of stochastic mapping methods, we also avoid the possible limitations of parsimony when it comes to character mapping — morphological evolution is often not parsimonious, at least for the traits we can observe back in time or study in detail today.

Fig. 3 – ML trait mapping on the tree in Fig. 2 (ie. considering molecular branch lengths). Note, the reconstruction of character state for the all-ancestor are ambiguous due to the extreme genetic distance between Fagus and the remainder of the Fagaceae. The situation in the scored fossils (Wilf et al. 2019, Denk et al. 2019) are shown for comparison.

For the ML mapping above, I scored intra-generic variations as additional states (ML ancestral-state reconstruction as implemented in the Mesquite program needs defined tips) and applied Mesquite's default model — this is essentially Lewis' Mk model for multi-state standard characters: one substitution category for any possible mutation. We can now compare the two mappings.

What our morphology-based tree recognized as derived was actually partly primitive. The near-synapomorphy (hemispheric indehiscent cupules) is in fact a symplesiomorphy of all Castanoideae + Quercus. Traits shared by Castanea-Castanopsis (pro parte, ie. some species show the ancestral, others the derived state) and Quercus are primitive, while those unique to (or part of intra-generic variation) one or several Castanoideae are derived.

Note that the alleged crown-group but old fossil Castanopsis rothwellii would fit at the base of the (core Fagaceae) tree (zero conflict) as well as close to its leaves (at least one conflicting character). Six of the seven traits can be pinpointed for the core Fagaceae ancestor. According to the reconstruction, it had three styles, scaly cupule appendages, hemispheric indehiscent cupules (vs. valvate in C. rothwellii), one flower per cupules, no valve dehiscence ("partial" in C. rothwellii), and inflorescences were unisexual and mixed (Wilf et al. state the Eocene fossils were unisexual, although the difference can only be assessed when investigating all inflorescences on a tree, see Denk et al.'s comment). The reconstruction is ambiguous regarding whether female flowers were clustered or solitary.

However, there is one implicit assumption held in common by all of the methods, including DNA-scaffolding, probabilistic and stochastic character mapping, total evidence dating, evolutionary placement algorithm (EPA) as implemented in the RAxML program, etc. That is: the inferred molecular tree is the true tree. This is the second fundamental flaw of cladistic approaches to evolution, as I will show in Part 2.

Data information

The morphological data used here is based on an emeneded version of the Wilf et al. matrix provided by my former colleague and co-author Thomas Denk (see also Denk et al. 2019, table 1); and it can be, together with the molecular data matrix used here, accessed via figshare.

References

Denk T, Grimm GW. (2009) The biogeographic history of beech trees. Review of Palaeobotany and Palynology 158: 83–100.

Denk T, Hill RS, Simeone MC, Cannon C, Dettmann ME, Manos PS. (2019) Comment on “Eocene Fagaceae from Patagonia and Gondwanan legacy in Asian rainforests”. Science 366: eaaz2189.

Grímsson F, Zetter R, Grimm GW, Krarup Pedersen G, Pedersen AK, Denk T. (2015) Fagaceae pollen from the early Cenozoic of West Greenland: revisiting Engler's and Chaney's Arcto-Tertiary hypotheses. Plant Systematics and Evolution 301: 809–832.

Hipp AL, Manos PS, Hahn M, et al. (2019) Genomic landscape of the global oak phylogeny. New Phytologist doi:10.1111/nph.16162.

Oh S-H, Manos PS. (2008) Molecular phylogenetics and cupule evolution in Fagaceae as inferred from nuclear CRABS CLAW sequences. Taxon 57: 434–451.

Renner SS, Grimm GW, Kapli P, Denk T. (2016) Species relationships and divergence times in beeches: New insights from the inclusion of 53 young and old fossils in a birth-death clock model. Philosophical Transactions of the Royal Society B doi:10.1098/rstb.2015.0135.

Wilf P, Nixon KC, Gandolfo MA, Cúneo NR. (2019) Eocene Fagaceae from Patagonia and Gondwanan legacy in Asian rainforests. Science 364:  eaaw5139.

For more literature, see the post:
Ockham's Razor applied but not used: can we make a DNA-scaffolding with seven characters?

Monday, November 12, 2018

More heretic bits: networks for (more) recent matrices published in Cladistics


This is Part 2 of a 2-part blog series. Part 1 covered some history, while this post has three (more) recently published matrices, and the take-home message.

Jumping forward in time, welcome to the 21st century

In Part 1, I showed several networks generated based on some early phylogenetic matrices published in the first volumes of the journal Cladistics. In this post, we will look at the most recent data matrices and trees uploaded to TreeBASE, covering the past seven years.

Nearly a generation later, and facing the "molecular revolution", some researchers (fortunately) still compile morphological matrices. This is an often overlooked but important work: genes and genomes can be sequenced by machines, and the only thing we need to do is to feed these machine-generated data into other powerful machines (and programs) to get a phylogenetic tree, or network. But no software and computer cluster can (so far) study anatomy, and generate a morphological matrix. The latter is paramount when we want to put fossils, usually devoid of DNA, in a (molecular) phylogenetic context. We need to do this when we aim to reconstruct histories in space and time.

Nevertheless, we can't ignore the fact that these important data are (still) far from tree-like. What holds for the matrices of the 80's (see the end of Part 1), still applies now.

So, let's have a look at the three most recent data sets (one morphological, two molecular) published in Cladistics that have their data matrix in TreeBASE.

The morphological dataset

Beutel et al. (2011; submission S11976) provided a "robust phylogeny of ... Holometabola", and note in their abstract: "Our results show little congruence with studies based on rRNA, but confirm most clades retrieved in a recent study based on nuclear genes."

Without having read the study, I can guess which clades (likely used here as a synonym for monophyletic group; but see David's post on Hennig and Cladistics) were confirmed. The data matrix contains: 356 multistate, with up to six states, characters scored and annotated for 34 taxa, including polymorphisms and some gaps ("–") viz missing data ("?"). Just by looking at the Neighbor-net inferred from this matrix. (Standard tree- or network-inference doesn't differ between gaps and missing data, but some people find it important to distinguish between "not applicable" and "not known" in a matrix.)

Neighbor-net inferred from simple pairwise distances computed based on Beutel et al.'s matrix. Brackets show my ad hoc assessment of candidates for monophyla (here: likely represented by clades in no matter how optimized trees).

How did I postulate the monophyla? By deduction: if two or more OTUs are much more similar to each other than to anything else in the matrix, they likely are part of the same evolutionary lineage, ie. have a common origin (= monophyletic in a pre-Hennigian sense). This, when the matrix well covers the group and morphospace, has a good chance to be inclusive (= monophyletic fide Hennig; for the covered OTUs). This is especially so when there is a good deal of homoplasy — the provided tree has a CI of 0.44 and RC of 0.33: convergences should be more randomly distributed than lineage-specific/-conserved traits. The latter don't need to be (or were, at some point in time) synapomorphies, shared derived unique traits, but could be diagnostic suites of characters that evolved in parallel within a lineage and passed on to all (or most) of the descendants.

The first molecular dataset

Let's look at the signal in the two molecular matrices.

In 2016, Gaspar and Almeida (submission S19167) tested generic circumscriptions in a group of ferns by "assembl[ing] the broadest dataset thus far, from three plastid regions (rbcL, rps4-trnS, trnL-trnF) ... includ[ing] 158 taxa and 178 newly generated sequences". They found: "three subfamilies each corresponding to a highly supported clade across all analyses (maximum parsimony, Bayesian inference, and maximum likelihood)."

The total matrix has 3250 characters, of which 1641 are constant and 1189 are parsimony-informative. This is a quite a lot for such a matrix, and, by itself, rules out parsimony for tree-inference. If half of the nucleotide sites are variable, then the rate of character change was high, and parsimony is statistically only robust, when the rate of change was low. High mutation rates or high level of divergence may also pose problems for distance methods and other optimality criteria, all closely related to parsimony.

The file includes three trees, labelled "vero" (which, in Italian, means "true"), "Fig._1" and "MPT". "Vero" and "Fig._1" come with branch lengths; judging from the values (<< 1), they are probabilistic trees (of some sort); the "MPT" is (as usual) provided as a cladogram without branch-lengths. It may be that the authors had to add the parsimony tree just to fulfill editorial policies, while being convinced "vero" is the much better tree. "Vero" is a fully resolved tree (the ML tree?), while "Fig._1" (Bayesian?) and "MPT" include polytomies.

Using PAUP*'s "describe" function, we learn that the "MPT" is 5101 steps long and has a CI of 0.41 and RC of 0.33. Nucleotide sequence data can be notoriously homoplasious, as we repeat the same four states into infinity and have to deal with an unknown but usually significant amount of back mutations. This adds to the other problems for parsimony:
  • transitions are more likely to happen than transversions; and
  • in coding gene regions, such as the rbcL, some sites (3rd codon positions) mutate much faster than others.
Still, parsimony trees are not necessarily wrong. Neither are NJ trees; and there are also datasets where probabilistic methods struggle, eg. when the likelihood surface of the treespace is flat.

So, the first question is: how different are the three trees provided? Rather than having to show three graphs, we can show the (strict) Consensus network of those trees.

A strict consensus network summarizing the topologies of the three trees provided in the TreeBASE submission of

The main difference is between "vero" and the other two — "Fig. 1" and the "MPT" are very similar (and both include polytomies). There are three main scenarios for a Consensus network like this with respect to the high portion of variable sites:
  1. "Fig. 1" is a Jukes-Cantor model-based tree,
  2. "Fig. 1" is an uncorrected p-distance based tree, or
  3. most of the variation is between ingroup (the subtree including all Blechnum) and outgroup (the other subtree).
"Vero" is still quite congruent, so the model used here can't be too much different, either.

What should ring one's alarm bells are, however, the many grade-like / staircase subtrees, which are unusual for a molecular data set. Staircases imply that each subsequent dichotomous speciation event resulted in a single species and a further diversifying lineage: multiple, consistently occurring budding events.

The same graph, with arrows showing grade evolution. Often found in morpho-data-based trees with ancestral, more ancient, and derived (from them), modern forms, but should ring an alarm bell when common in a molecular tree. Major clades (found in all three trees) are labelled for comparison with the next graph.

Let's compare this to the Neighbor-net (usually, I would use model-based distances in such a case, but here we can do with uncorrected p-distances).

A Neighbor-net inferred from uncorrected p-distances based on Gaspar & Almeida's matrix; the major clades are labelled as in the preceding graph. Note the isolated, long-branch blue dots with asterisks, indicating the position of the first diverged species in the large clades G and I. Genuine signal or missing data artefact?

The Neighbor-net shows only a limited number of tree-like portions, but does correspond with the main clades above. Only A and B are dissolved, which are the two first diverging clades in the original trees (preceding graph). Some OTUs are placed close to the centre of the graph, or even along a tree-like portion (purple dots), a behaviour known from actual ancestors: some OTUs apparently have sequences that may be literally ancestral to others. This explains the grade structure seen in the original trees. Others (violet dots) create boxes, which may reflect a genuine ambiguous signal, or just be missing data leading to ambiguous pairwise distances. The latter (missing data artefact) is behind the misplacement of the four OTUs (red dots): missing data can inflate pairwise distances severely. And, like parsimony, distance-based methods are more vulnerable to long-branch(edge)-attraction than probabilistic methods.

Model-based distances may help clean up this a bit, but the networks needed for these kind of data are Support consensus networks (see e.g. Schliep et al., MEE, 2017). The split appearance of the Neighbor-net hints at internal signal conflict and, with respect to the high number of variable sites (note the sometimes extremely long terminal edges), saturation issues. Two major questions would be:
  1. How do the different markers (coding gene vs. inter-genic spacers with different levels of diversity; rps4-trnS is typically more divergent than the trnL-trnF spacer) resolve relationships, which clades / topological alternatives receive unanimous support?
  2. Does it make a difference to run a fully partitioned (ML) analysis vs. an unpartitioned one vs. one excluding the 3rd codon position in the gene?
For intra-clade evolutionary pathways, it would be worthwhile to give median networks and suchlike a try, as parsimony methods that can discern ancestor-descendant relationships.

The second molecular dataset

The most recent data are from Kuo et al. (2017; submission S20277), who inferred a "robust ... phylogeny" (see Part 1, Jamieson et al. 1987, and Beutel et al., above) for a group of ferns, focusing on the taxonomy of a single genus, Deparia, that now includes five traditionally recognized genera. In the abstract it says: "... seven major clades were identified, and most of them were characterized by inferring synapomorphies using 14 morphological characters".

The matrix includes the molecular characters used to infer the major clades plus two trees, labelled "bestREP1" and "rep9BEST", both with branch lengths. Branch length values indicate that "bestREP1" could be parsimony-optimized (with averaged or weighted branch lengths), while "rep9BEST" is either a ML or Bayesian tree (technically, it could be a distance-based tree, too, but I don't think such "phenetics" are condoned by Cladistics).

Re-calculated, the first tree ("bestREP1") is shorter (3024 steps) than the one of Gaspar & Almeida, reflecting the much lower number of parsimony-informative sites (979). Many of the sites differ only between the focal genus and the outgroups, which is well visible in the Neighbor-net. [For those of you unfamiliar with Neighbor-nets, a parsimony analysis of these data takes hours, or days depending on the software and computer, while the distance matrix and the resultant Neighbor-net is inferred in a blink.]

The Neighbor-net based on Kuo et al.'s data. Why do we need to include long-branching, distant outgroups when we just want to bring order in a genus? Because to test monophyly, we need a rooted tree (ambiguous or not, or even biased by branching artefacts).

Let's remove the distant, long-branching outgroups, which (as we can see in the Neighbor-net) at best provide ambiguous signal for rooting the ingroup — at worst, they trigger ingroup-outgroup branching artefacts. What could a Neighbour-net have contributed regarding taxonomy and the seven major monophyletic intrageneric groups ("clades")? Pretty much everything needed for the paper, I guess (judging from the abstract).

Same data as above, but outgroups removed. The structure of this Neighbour-net allows to identify seven likely candidates for monophyla ("1"–"7"), with "1" and "2" being obvious sister lineages. Colours refer to the clusters ("A"–"E") annotated above.

On a side note: by removing the long-branching, distant outgroups, taxon "T" is resolved as a probable member of the putative monophyletic group "5" (= "E" in the full graph with outgroups, and surely a high-supported subtree in any ingroup-only reconstruction, method-independent). Placing the root between "T" and the rest of the genus implies that "5" is a paraphyletic group comprising species that haven't evolved and diversified at all (ie. are genetically primitive), in stark contrast to the other main intra-generic lineages. This is not impossible, but quite unlikely. More likely is the second scenario (primary split between "1"–"3" and "4"–"7"). Having "4" as sister to the rest could be an alternative, too.

This is where Hennig's logic could be of help: find and tabulate putative synapomorphies to argue for a set and root that makes the most sense regarding morphological evolution and molecular differentiation.

The take-home message(s)

We have argued before that it is in the ultimate interest of science and scientists to give access to phylogenetic data. No matter where one stands regarding phylogenetic philosophy, we should publish our data, so that people can do analyses of their own. Discussion should be based on results, not philosophies.

When you deal with morphological data, you should never be content with inferring a single tree (parsimony or other). You have to use networks.

The Neighbor-net was born as late as 2002 (Bryant & Moulton, 2002, in: Guigó R, and Gusfield D, eds, Algorithms in Bioinformatics, Second International Workshop, WABI, p. 375–391; paywalled) and made known to biologists in 2004 (same authors, same title, in Mol. Biol. Evol. 21:255–265), so that authors before this time did not have access to its benefits. Similarly, Consensus networks arrived around about the same time (Holland & Moulton 2003, in: Benson G, and Page R, eds, Algorithms in Bioinformatics: Third International Workshop, WABI, p. 165–176). However, the Genealogical World of Phylogenetic Networks has been here for six years now (first post February 2012). So there is now no excuse for publishing a cladogram without having explored the tree-likeness of your matrix' signal!

Neighbor-nets like the ones I showed in this 2-piece post (or can be found in many of our other posts) are a quick and essential tool to explore the basic signal in your matrix:
  • How tree-like is it?
  • Where are the potential conflicts, obscurities?
  • What are the principal evolutionary alternatives (competing topologies)?
  • What is well supported (especially regarding taxonomy and the question of monophyly)?
Even if you don't use it in your paper, the network will tell you what you are dealing with when you start inferring trees.

The second essential tool is the much under-used Support consensus network, not shown in this post but in plenty of our other posts (and many papers I co-authored; for a comprehensive collection of network-related literature see Who's who in phylogenetic networks by Philippe Gambette). Support consensus networks estimate and visualize the robustness of the signal for competing topological (tree) alternatives.

Consensus networks should also be obligatory for those molecular data,where even probabilistic methods fail to find a single fully resolved, highly supported tree.

If the editors of Cladistics are really dedicated to parsimony, they should not still insist only on a parsimony tree (often provided as cladogram), but also parsimony-based networks as well:
  • strict Consensus networks to summarize the MPT samples instead of the standard strict Consensus cladograms;
  • bootstrap Support consensus networks showing the signal strength and support for alternative trees/competing clades (TNT has many bootstrapping options to play around with); and
  • Median networks and such-like for datasets with few mutations, and low levels of expected homoplasy.
This is what the 2016 #parsimonygate uproar (see Part 1) should have been about (12 years after Neighbor-nets, and 11 years after Consensus networks). Not the prioritizing of parsimony, but the naivety or ignorance towards pitfalls of (parsimony or other) trees inferred from data not providing tree-like signal or riddled by internal conflict.
This is a problem not limited to Cladistics, but found, to my modest experience in professional science (c. 20 years), in many other journals as well (e.g. Bot. J. Linn. Soc., Taxon, Mol. Phyl. Evol., J. Biogeogr., Syst. Biol., Nature, Science).

Hence, here are my suggestions for future conference buttons, instead of those shown in Part 1.

No Cladograms! Use Neighbour-nets! Support Consensus Networks as obligatory!

Further reading for those who mistrust trees or become network-curious in general

Monday, November 5, 2018

A bit of heresy: networks for matrices used in Cladistics studies


[This is Part 1 of a two-part topic – this one is Historical matrices from the 1980s]

When I first came into contact with phylogenetics (usually based on morphological data sets, back then) and after reading Hennig's book (the original German version, published in 1950), I dreamed about publishing in Cladistics, the journal of the Willi Hennig Society (WHS). I never did. In this post, I show why.

Later on, in 2016, Cladistics achieved renewed fame due to an editorial that triggered a twitter uproar under the hashtag #parsimonygate. A lot of people were shocked to read in the editorial that the journal (still) prefers and requires parsimony-based inferences (in fact, parsimony-based trees). Some people, like Joe Felsenstein, were not at all surprised. I wasn't either, because Cladistics is the journal of the Willi Hennig Society (WHS), which has always been dedicated to parsimony: "Ockham told Popper told Hennig to use parsimony" (see the historical summary by Felsenstein in Systematic Biology, 2001; free access).

Historical buttons that you (allegedly) could get at meetings of the WHS. Left: Joe Felsenstein; right: L for Likelihood. Just a gag, of course! Nothing serious behind it.

In the good old days, when the "Phylogenetic Wars" were still on (in the 1980s, petering out in the 90s), they would invite a probability-ist to their conference to tear him down. My first phylogenetic paper (2002) got a negative review (ie. rejection, invitation to resubmit) by a WHS member solely because it did not include a parsimony tree, which he described as "standard these days". More recently, they ensured free access to TNT, the current main software for doing parsimony analysis and an essential tool for many palaeontologists.

I stopped using parsimony trees very early in my career, but I'm still a great fan of the family of methods based on median networks, which operate under the same parsimony criterion (Clades, Cladograms, ...; Using Median networks ...). Fate exposed me early to the Neighbor-nets, which can be used as a quick check of how tree-like the signal is in data matrices, to start with.

The thing that bugged me most concerning many journals, including Cladistics, is not a focus on parsimony, but the lack of data documentation and easy data access. To me, it seems natural to use a service like TreeBASE, when my main dedication is to tree-inference. TreeBASE allows you to provide your data and inferred trees to the general public in the common NEXUS format, so that other people can make use of it.

Luckily, some authors of Cladistics upload their data (about one study per 1–3 years). So, here are some data-display networks showing the strengths and weaknesses of the parsimony trees in the original publications, which have been randomly selected from among the oldest ones and the newest ones (I found) in TreeBASE. I won't discuss the actual results, as Cladistics is pay-walled, so just enjoy the graphs.

The oldest one (in my list), Dahlgren & Bremer 1985, TreeBASE submission number S231

The submission (a binary matrix, including some missing data; published in the first volume of Cladistics) comes with three angiosperm trees: one composite order-level tree, plus two empirical trees labelled as "Fig. 2" and "Fig. 3" using the family-level OTUs in the matrix. The latter two look like this:

Connected cladograms of "Fig. 2" and "Fig. 3", the result of two parsimony analyses. Jumping taxa/clades highlighted with colours.
That the matrix is not only highly homoplasious (CI = 0.28) but has a severe signal problem, becomes obvious when inferring a NJ tree, providing a third topology.

A NJ tree (fulfilling least-squares optimality criterion for phylogenetic trees) from the same matrix: blue, branches incongruent among the original trees and the NJ tree. Color coding: light blue, branch congruent to "Fig. 2" tree (different in "Fig. 3" tree); green, branch found in all three trees; red, branch incongruent to consistent placement in both original trees.

Not surprisingly, the Neighbor-net inferred from simple (mean) Hamming distances is a spider-web, as the matrix' signal is not tree-like at all — all non-green branches above, or their conflicting alternatives, receive low to very low bootstrap support, independent of the optimality criterion used.

The Neighbor-net inferred from Dahlgren & Bremer's matrix.

Despite its spider-web structure, we do learn quite a lot from the Neighbor-net regarding what is behind the clades in the original trees. For example, we can overlay a Dahlgrenogram representing the top-most subtree of the "Fig. 2" tree.

Blue, red and yellow fields denote (sub)clades in Dahlgren & Bremer's "Fig. 2" tree that compose the top clade (grey).

The same could be done for all the other clades.

TreeBASE submission S329, worms (Oligochaeta) by Jamieson et al. (1987)

The more perfect is a character matrix regarding tree-inference (ie. with tree-compatible characters), the more similar the NJ and the parsimony-tree will be (or any other tree, under any other optimality criterion), as we can see in this second example published in the third volume of Cladistics.

The tree (the abstract notes a single most-parsimonious tree) was inferred from a multistate matrix with up to seven states, possibly including some characters that should be treated as ordered, but such specifics are not included in the original NEXUS file, so we will treat them as unordered.

Aside from grades becoming clades (and vice versa), the published tree (unordered: 102 steps, high CI = 0.81, RC = 0.53) and the NJ tree are quite similar, even regarding their relative branch-lengths.

Two phylograms: left, the original MPT, right, a NJ tree, shared branches in green, (partly) conflicting ones in orange. Cladists address the left tree as "phylogenetic", the right one as "phenetic", but both are equally valid solutions using different optimality criteria.

Moreover, the Neighbor-net is much less complex than in the previous examples, with individual edges corresponding to branches in both trees — Neighbor-nets are truly meta-phylogenetic graphs.

Splits found in the original MPT in green, when corresponding with edges in the Neighbour-net, and orange, when there is no corresponding edge (according to the abstract, the authors discuss alternatives to certain branches in their tree). Edges found in the NJ tree (providing an alternative topology/phylogenetic hypothesis) in blue.

Submission S349, an amniote phylogeny by Gaulthier et al. (1988)

This is a matrix much to my liking, as it includes extinct taxa, with quite impressive dimensions (computers back in 1988 were awfully slow): 316 characters with up to four states for 31 taxa. Naturally, it includes a lot of missing data, as do all fossil-including matrices.

Missing data is potentially a bigger problem for distance-based approaches than for character-based ones like parsimony, maximum likelihood or Bayesian inference — when there is little character overlap between the fossil taxa, their pairwise distances will be distorted. Missing data can be an equal problem for tree-inference — depending which characters are missing, many different topologies are equally optimal, or nearly so. In Gaulthier et al.'s matrix 10% of the characters are parsimony-uninformative.

Similar to the angiosperm matrix, Gaulthier et al.'s tree has a relatively low CI (0.45) and RC (0.33), i.e. there is homoplasy adding to the missing data as a source of incompatible, tree-unlike signals.

Just by comparing the NJ tree to the parsimony tree, we can see that distance distortion because of missing data is no big deal for this matrix.


The trees are largely congruent, with three striking exceptions: the birds (Aves), the crocodiles (Crocodylia) and turtles (Testudines) are not placed as sisters to the lineage leading to modern-day mammals (tree provided by Gaulthier et al.), but fall in the "dinosaur"-only clade in the NJ tree (compare with the current Tree of Life: Archosauria). This makes sense (data-wise), because in Gaulthier's matrix the taxon pairs Aves + Ornithosuchia and Crocodylia + Pseudosuchia are identical in their shared defined characters (ie. zero-distance pairs). Obviously, the parsimony tree comes with some implicit assumptions: the unweighted/unordered single most-parsimonious tree PAUP* infers for the matrix using the branch-and-bound algorithm has only 510 steps, a higher CI (0.66) and RC (0.59), and is largely congruent with the NJ tree; except that Captorhinidae and Testudines are sisters and Casea, Ophiacodon and Edaphosaurus form a grade not a clade.

As in the other cases so far, the Neighbor-net well captures the actual data situation.

Blue edge bundles refer to splits shared with both the NJ tree and the (inferred, not provided) MPT. Note that some splits in the NJ tree and or the MPT have no counterpart in the Neighbour-net. One split found in the MPT but not in the NJ tree has a corresponding edge in the Neighbour-net (light blue).
The thin "upper trunk" in the Neighbor-net further shows that the matrix provides a strong signal for an increase of shared derived ('mammalian') and decrease of shared ancestral ('reptilian') traits, which is a bias. Although the MPT and NJ tree agree well, the matrix provides clear tree-like signal only for terminal relationships in the other main, inferred clade. The thinning trunk may also indicate a taxon sampling issue. Well-sampled phylogenetic data sets usually result in more star-like networks (see eg. graphs in this post on fossil and extant walnuts, dinosaurs, spermatophytes, or the above ones and the next one) in contrast to non-phylogenetic data sets (see eg. the posts on breast sizes, airlines, or moons)


Take-home message in the middle of the film

Even though they are arbitrary choices, the three matrices above show what phylogeneticists had to work with in the 1980s morphological datasets:
  • ... trapped in homoplasy (Dahlgren & Bremer, 1985) — datasets in which phylogenetic relationships were obscured behind highly ambiguous, non-treelike signal;
  • ... asking for a model (Jamieson et al., 1987) — datasets with partly consistent signal, but not consistent enough to result in the same tree independent of the optimality criterion;
  • ... encoding a tree (Gaulthier et al., 1988) — datasets tweaked to promote a certain evolutionary hypothesis, including (superficially) simple series of gradual evolution and ancestor-descendant pairs (see Trivial data, not so trivial graphs). Such data will result in a single optimal tree (method independent!) dominanted by staircase-like subtrees. This may be fine for a cladist, but nothing a phylogeneticist / evolutionary biologist could really be content with (not in the 1980s, or before 1950).


Top, two phylogenetic tress sketched by Darwin; bottom, Hilgendorf's (1866) phylogenetic tree. There are quite a few before 1950 (eg. Pojárkova, 1933, Acta Institute of Botany, Academy of Sciences of the USSR, ser. 1, 1: 225–374; unfortunately have no copy/scan)

Tuesday, October 24, 2017

Let's distinguish between Hennig and Cladistics


There are theoretically an infinite number of ways to mathematically analyze any set of data, and yet it is unlikely that all (or even most) of these will have any relevance to a study of biology. In this sense, the philosophy of phylogenetic analysis needs to show that there is a strong basis for treating any particular mathematical analysis as having biological relevance. This is a point that I have discussed before: Is there a philosophy of phylogenetic networks?

Willi Hennig clearly has some role to play here. However, his ideas are often treated as being solely related to one particular form of phylogenetic analysis — cladistics. In this post I will point out that his work has a much greater relevance than that — he provides a crucial logical step that applies to all phylogenetic inference.

The steps of phylogenetic inference are shown in the first figure, which is taken from my earlier post. The first step is a mathematical inference from character data to tree/network; the second step is a logical inference that the mathematical summary resulting from the first step has some biological relevance; and the third step is a practical inference that the biological summary applies to whole organisms as well as to their characters.

The logic of phylogeny reconstruction

Summary

Hennig's concept of "shared innovations" (which he called synapomorphies) is the only thing that allows us to use the mathematical phylogenetics in the pursuit of genealogical history. Without this concept, the mathematics could just produce something like the arithmetic mean, a mathematical concept with no connection to real objects (unlike the median or mode, which will always be real). The idea of shared innovations is what leads us to believe that the mathematical summary (whether tree or network) might actually also be a close approximation to the real thing. This is a separate concept from cladistics, which is simply a mathematical algorithm based on a particular optimality criterion (parsimony), just like maximum likelihood or bayesian approaches. So, shared innovations underlie the use of both parsimony, likelihood and distance methods — Willi Hennig (and, before him, Karl Brugmann in linguistics) is relevant no matter what algorithm we use.

Mathematical analyses

If they are to represent genealogical history, then all trees and networks in phylogenetics will be directed acyclic graphs (DAGs), mathematically. There are many ways to produce a DAG, some of which have had varying degrees of popularity in phylogenetics, and some of which have not been used at all.

To produce an acyclic line graph (in which nodes are connected by edges), we can start with character data or distance data. We can then use various optimality criteria to choose among the many graphs that could apply to the data, such as parsimony (usually ssociated with cladistics) and likelihood (either as maximum likelihood or integrated likelihood). We can also ensure that the graph is directed (ie. the edges have arrows), by choosing a root location, either directly as part of the analysis or a posteriori by specifying an outgroup.

All of these approaches are mathematically valid, as are a number of others. They all provide a mathematical summary of the data. This is step one of the phylogenetic inference, as illustrated above.

But what of step two? Biologists need a summary of the data that has biological relevance, as well, not just mathematical relevance. This has long been a thorn in the side of biologists — just because they can perform a particular mathematical calculation does not automatically mean that the calculation is relevant to their biological goal.

Consider the simplest mathematics of all — calculating the central location of a set of data. There are many ways to do this, mathematically — indeed, there are technically an infinite number of ways. These include the mode, the median, the arithmetic mean, the geometric mean, and the harmonic mean. All of these are mathematically valid, but do any of them produce a central location that describes biology?

The mode does, because it is the most common observation in the dataset. The median usually does, because it is the "middle" observation in the dataset. But what of the various means? There is no necessary reason for them to describe biology, although they are perfectly valid mathematics.

For instance, the modal number of children in modern families is 2, meaning that more families have this number than any other number of children. The median number is also 2, meaning that half of the families have 2 or fewer children and half of the families have 2 or more. So, these mathematical summaries are also descriptions of real families. But the means are not. For example, the arithmetic mean number of children is 2.2, which does not describe any real family. If you ever find a family with 2.2 children, then you should probably call the police, to investigate!

Mathematically valid data summaries have a lot of relevance, but they do not necessarily describe biological concepts. I can use the mean number of children per local family to estimate the number of schools that I might need in that area, but I cannot use it to describe the families themselves. This is a classic case of "horses for courses".

Hennig

So, in phylogenetics we need some piece of logic that says that we can expect our DAG (a mathematical concept) to be a representation of a genealogy (a biological concept). Our genealogical estimate may still be wrong (and indeed it probably will be, in some way!), but that is a separate issue. The DAG needs to a reasonable representation, not a correct one. Correctness needs to be a result of our data, not our mathematics.

This is where Willi Hennig comes in. Hennig's ideas, and the ideas derived from them, are illustrated in the second figure.


Hennig explicitly noted that characters have a genealogical polarity, with ancestral states being modified into derived states through evolutionary time. Furthermore, he noted that it is only the derived states that are of relevance to studying evolutionary history — the sharing of derived character states reveals evolutionary history, but shared ancestral states tells us nothing.

We have done two things with these Hennigian ideas. Some people have been interested in classification, for which the concept of monophyly is relevant, and others have been interested in reconstructing the genealogies, rather than simply interpreting them.

Phylogenetics

Reconstructing a tree-like phylogenetic history is conceptually straightforward, although it took a long time for someone (Hennig 1966) to explain the most appropriate approach. Interestingly, the study of historical linguistics has developed the same methodology (Platnick and Cameron 1977; Atkinson and Gray 2005), thus independently arriving at exactly the same solution to what is, in effect, exactly the same problem. From this point of view, the logical inference itself is uncontroversial; and its generic nature means that it can be used for any objects with characteristics that can be identified and measured, and that follow a history of descent with modification. I will, however, discuss this in terms of biology — you can make the leap to other objects yourself.

The objective is to infer the ancestors of the contemporary organisms, and the ancestors of those ancestors, etc., all the way back to the most recent common ancestor of the group of organisms being studied. Ancestors can be inferred because the organisms share unique characteristics (shared innovations, or shared derived character states. That is, they have features that they hold in common and that are not possessed by any other organisms. The simplest explanation for this observation is that the features are shared because they were inherited from an ancestor. The ancestor acquired a set of heritable (i.e. genetically controlled) characteristics, and passed those characteristics on to its offspring. We observe the offspring, note their shared characteristics, and thus infer the existence of the unobserved ancestor(s). If we collect a number of such observations, what we often find is that they form a set of nested groupings of the organisms.

Hennig, in particular, was interested in the interpretation of phylogenetic trees, rather than their reconstruction. He did this interpretation in terms of monophyletic groups (also called clades), each of which consists of an ancestor and all of its descendants. These are natural groups in terms of their evolutionary history, whereas other types of groups (eg. paraphyletic, polyphyletic) are not. So, a phylogenetic tree consists of a set of nested clades, which are the groups that are represented and given names in formal taxonomic schemes.

For phylogenetic trees, there is thus a rationale for treating a tree diagram as a representation of evolutionary history. For example, in a study of a set of gene sequences, first we produce a mathematical summary of the data based on a quantitative model. We then infer that this summary represents the gene history, based on the Hennigian logic that the patterns are formed from a nested series of shared innovations (this is a logical inference about the biology being represented by the mathematical summary). We then infer that this gene history represents the organismal history, based on the practical observation that gene changes usually track changes in the organisms in which they occur (ie. a pragmatic inference).

Mis-interpretations of Hennig

What I have said above has lead to various mis-interpretations of Hennig's role in phylogenetics.

First, he did not propose any specific method for producing a phylogenetic tree (or network). He was concerned about the logic of the diagram. not how to get it in the first place. He distinguished shared derived character states, or shard innovations, (he called them synapomorphies) from shared ancestral states (symplesiomorphies), and noted that only the former are relevant for phylogenies. So, distance methods will also work in phylogenetics provided the distances are based on homologous apomorphic features. If they are not so based, then they are simply mathematical constructions, which may or may not represent anything to do with phylogeny. Distances estimated from plesiomorphic features can be used to construct a tree, obviously, but there is no reason to expect that tree to represent a phylogeny.

Second, parsimony analysis was developed independently of Hennig, by people such as Farris, Nelson and Platnick. This came to be called cladistics, intended by Ernst Mayr to be a derogatory term for the new form of analysis. The fact that the Willi Hennig Society is associated exclusively with cladistics has nothing to do with Hennig himself, or with the logic of his approach to phylogenetics. You need to clearly distinguish between Hennig and Cladistics!

Third, Hennig was more interested in classification than he was in phylogeny reconstruction. This seems to cause confusion for gene jockeys and linguists, in particular, who often associate phylogenetics solely with classification (see, for example, Felsenstein 2004, chapter 10). Sure, Hennig was primarily interested in the interpretation of phylogenies, rather than their construction. However, that was simply a personal point of view. The logic of his work transcends his own personal interests. Without him, no genealogical reconstruction makes logical sense, in genetics or linguistics. Mathematical methods for summarizing data were developed independently in genetics and linguistics, just as they were in other areas of biology and also in stemmatology. However, without the concept of shared innovations, these methods remain mathematical summaries, not estimates of genealogies.

Finally, Hennig's work was not original, being naturally a synthesis of much previous work. In biology, the work of Walter Zimmerman is frequently noted (eg. Donoghue & Kadereit 1992), and in linguistics the work of Karl Brugmann is obviously important (see Mattis' post Arguments from authority, and the Cladistic Ghost, in historical linguistics). Sometimes, wheels have to be re-invented many times before the general populace comes to realize just how important they are.

References

Atkinson QD, Gray RD (2005) Curious parallels and curious connections — phylogenetic thinking in biology and historical linguistics. Systematic Biology 54: 513-526.

Donoghue MJ, Kadereit W (1992) Walter Zimmermann and the growth of phylogenetic theory. Systematic Biology 41: 74-85.

Felsenstain J (2004) Inferring Phylogenies. Sinauer Associates, Sunderland MA.

Hennig W (1966) Phylogenetic Systematics. University of Illinois Press, Urbana IL. [Translated by DD Davis and R Zangerl from W. Hennig 1950. Grundzüge einer Theorie der Phylogenetischen Systematik. Deutscher Zentralverlag, Berlin.]

Platnick NI, Cameron HD (1977) Cladistic methods in textual, linguistic, and phylogenetic analysis. Systematic Zoology 26: 380-385.

Tuesday, September 19, 2017

Arguments from authority, and the Cladistic Ghost, in historical linguistics


Arguments from authority play an important role in our daily lives and our societies. In political discussions, we often point to the opinion of trusted authorities if we do not know enough about the matter at hand. In medicine, favorable opinions by respected authorities function as one of four levels of evidence (admittedly, the lowest) to judge the strength of a medicament. In advertising, the (at times doubtful) authority of celebrities is used to convince us that a certain product will change our lives.

Arguments from authority are useful, since they allow us to have an opinion without fully understanding it. Given the ever-increasing complexity of the world in which we live, we could not do without them. We need to build on the opinions and conclusions of others in order to construct our personal little realm of convictions and insights. This is specifically important for scientific research, since it is based on a huge network of trust in the correctness of previous studies which no single researcher could check in a lifetime.

Arguments from authority are, however, also dangerous if we blindly trust them without critical evaluation. To err is human, and there is no guarantee that the analysis of our favorite authorities is always error proof. For example, famous linguists, such as Ferdinand de Saussure (1857-1913) or Antoine Meillet (1866-1936), revolutionized the field of historical linguistics, and their theories had a huge impact on the way we compare languages today. Nevertheless, this does not mean that they were right in all their theories and analyses, and we should never trust any theory or methodological principle only because it was proposed by Meillet or Saussure.

Since people tend to avoid asking why their authority came to a certain conclusion, arguments of authority can be easily abused. In the extreme, this may accumulate in totalitarian societies, or societies ruled by religious fanatism. To a smaller degree, we can also find this totalitarian attitude in science, where researchers may end up blindly trusting the theory of a certain authority without further critically investigating it.

The comparative method

The authority in this context does not necessarily need to be a real person, it can also be a theory or a certain methodology. The financial crisis from 2008 can be taken as an example of a methodology, namely classical "economic forecasting", that turned out to be trusted much more than it deserved. In historical linguistics, we have a similar quasi-religious attitude towards our traditional comparative method (see Weiss 2014 for an overview), which we use in order to compare languages. This "method" is in fact no method at all, but rather a huge bunch of techniques by which linguists have been comparing and reconstructing languages during the past 200 years. These include the detection of cognate or "homologous" words across languages, and the inference of regular sound correspondence patterns (which I discussed in a blog from October last year), but also the reconstruction of sounds and words of ancestral languages not attested in written records, and the inference of the phylogeny of a given language family.

In all of these matters, the comparative method enjoys a quasi-religious authority in historical linguistics. Saying that they do not follow the comparative method in their work is among the worst things you can say to historical linguists. It hurts. We are conditioned from when we were small to feel this pain. This is all the more surprising, given that scholars rarely agree on the specifics of the methodology, as one can see from the table below, where I compare the key tasks that different authors attribute to the "method" in the literature. I think one can easily see that there is not much of an overlap, nor a pattern.

Varying accounts on the "comparative methods" in the linguistic literature

It is difficult to tell how this attitude evolved. The foundations of the comparative method go back to the early work of scholars in the 19th century, who managed to demonstrate the genealogical relationship of the Indo-European languages. Already in these early times, we can find hints regarding the "methodology" of "comparative grammar" (see for example Atkinson 1875), but judging from the literature I have read, it seems that it was not before the early 20th century that people began to introduce the techniques for historical language comparison as a methodological framework.

How this framework became the framework for language comparison, although it was never really established as such, is even less clear to me. At some point the linguistic world (which was always characterized by aggressive battles among colleagues, which were fought in the open in numerous publications) decided that the numerous techniques for historical language comparison which turned out to be the most successful ones up to that point are a specific method, and that this specific method was so extremely well established that no alternative approach could ever compete with it.

Biologists, who have experienced drastic methodological changes during the last decades, may wonder how scientists could believe that any practice, theory, or method is everlasting, untouchable and infallible. In fact, the comparative method in historical linguistics is always changing, since it is a label rather than a true framework with fixed rules. Our insights into various aspects of language change is constantly increasing, and as a result, the way we practice the comparative method is also improving. As a result, we keep using the same label, but the product we sell is different from the one we sold decades ago. Historical linguistics are, however, very conservative regarding the authorities they trust, and our field was always very skeptical regarding any new methodologies which were proposed.

Morris Swadesh (1909-1967), for example, proposed a quantitative approach to infer divergence dates of language pairs (Swadesh 1950 and later), which was immediately refuted, right after he proposed it (Hoijer 1956, Bergsland and Vogt 1962). Swadesh's idea to assume constant rates of lexical change was surely problematic, but his general idea of looking at lexical change from the perspective of a fixed set of meanings was very creative in that time, and it has given rise to many interesting investigations (see, among others, Haspelmath and Tadmor 2009). As a result, quantitative work was largely disregarded in the following decades. Not many people payed any attention to David Sankoff's (1969) PhD thesis, in which he tried to develop improved models of lexical change in order to infer language phylogenies, which is probably the reason why Sankoff later turned to biology, where his work received the appreciation it deserved.

Shared innovations

Since the beginning of the second millennium, quantitative studies have enjoyed a new popularity in historical linguistics, as can be seen in the numerous papers that have been devoted to automatically inferred phylogenies (see Gray and Atkinson 2003 and passim). The field has begun to accept these methods as additional tools to provide an understanding of how our languages evolved into their current shape. But scholars tend to contrast these new techniques sharply with the "classical approaches", namely the different modules of the comparative method. Many scholars also still assume that the only valid technique by which phylogenies (be it trees or networks) can be inferred is to identify shared innovations in the languages under investigation (Donohue et al. 2012, François 2014).

The idea of shared innovations was first proposed by Brugmann (1884), and has its direct counterpart in Hennig's (1950) framework of cladistics. In a later book of Brugmann, we find the following passage on shared innovations (or synapomorphies in Hennig's terminology):
The only thing that can shed light on the relation among the individual language branches [...] are the specific correspondences between two or more of them, the innovations, by which each time certain language branches have advanced in comparison with other branches in their development. (Brugmann 1967[1886]:24, my translation)
Unfortunately, not many people seem to have read Brugmann's original text in full. Brugmann says that subgrouping requires the identification of shared innovative traits (as opposed to shared retentions), but he remains skeptical about whether this can be done in a satisfying way, since we often do not know whether certain traits developed independently, were borrowed at later stages, or are simply being misidentified as being "shared". Brugmann's proposed solution to this is to claim that shared, potentially innovative traits, should be numerous enough to reduce the possibility of chance.

While biology has long since abandoned the cladistic idea, turning instead to quantitative (mostly stochastic) approaches in phylogenetic reconstruction, linguists are surprisingly stubborn in this regard. It is beyond question that those uniquely shared traits among languages that are unlikely to have evolved by chance or language contact are good proxies for subgrouping. But they are often very hard to identify, and this is probably also the reason why our understanding about the phylogeny of the Indo-European language family has not improved much during the past 100 years. In situations where we lack any striking evidence, quantitative approaches may as well be used to infer potentially innovated traits, and if we do a better job in listing these cases (current software, which was designed by biologists, is not really helpful in logging all decisions and inferences that were made by the algorithms), we could profit a lot when turning to computer-assisted frameworks in which experts thoroughly evaluate the inferences which were made by the automatic approaches in order to generate new hypotheses and improve our understanding of our language's past.

A further problem with cladistics is that scholars often use the term shared innovation for inferences, while the cladistic toolkit and the reason why Brugmann and Hennig thought that shared innovations are needed for subgrouping rests on the assumption that one knows the true evolutionary history (DeLaet 2005: 85). Since the true evolutionary history is a tree in the cladistic sense, an innovation can only be identified if one knows the tree. This means, however, that one cannot use the innovations to infer the tree (if it has to be known in advance). What scholars thus mean when talking about shared innovations in linguistics are potentially shared innovations, that is, characters, which are diagnostic of subgrouping.

Conclusions

Given how quickly science evolves and how non-permanent our knowledge and our methodologies are, I would never claim that the new quantitative approaches are the only way to deal with trees or networks in historical linguistics. The last word on this debate has not yet been spoken, and while I see many points critically, there are also many points for concrete improvement (List 2016). But I see very clearly that our tendency as historical linguists to take the comparative method as the only authoritative way to arrive at a valid subgrouping is not leading us anywhere.

Do computational approaches really switch off the light which illuminates classical historical linguistics?

In a recent review, Stefan Georg, an expert on Altaic languages, writes that the recent computational approaches to phylogenetic reconstruction in historical linguistics "switch out the light which has illuminated Indo-European linguistics for generations (by switching on some computers)", and that they "reduce this discipline to the pre-modern guesswork stage [...] in the belief that all that processing power can replace the available knowledge about these languages [...] and will produce ‘results’ which are worth the paper they are printed on" (Georg 2017: 372, footnote). It seems to me, that, if a discipline has been enlightened too much by its blind trust in authorities, it is not the worst idea to switch off the light once in a while.

References
  • Anttila, R. (1972): An introduction to historical and comparative linguistics. Macmillan: New York.
  • Atkinson, R. (1875): Comparative grammar of the Dravidian languages. Hermathena 2.3. 60-106.
  • Bergsland, K. and H. Vogt (1962): On the validity of glottochronology. Current Anthropology 3.2. 115-153.
  • Brugmann, K. (1884): Zur Frage nach den Verwandtschaftsverhältnissen der indogermanischen Sprachen [Questions regarding the closer relationship of the Indo-European languages]. Internationale Zeischrift für allgemeine Sprachewissenschaft 1. 228-256.
  • Bußmann, H. (2002): Lexikon der Sprachwissenschaft . Kröner: Stuttgart.
  • De Laet, J. (2005): Parsimony and the problem of inapplicables in sequence data. In: Albert, V. (ed.): Parsimony, phylogeny, and genomics. Oxford University Press: Oxford. 81-116.
  • Donohue, M., T. Denham, and S. Oppenheimer (2012): New methodologies for historical linguistics? Calibrating a lexicon-based methodology for diffusion vs. subgrouping. Diachronica 29.4. 505–522.
  • Fleischhauer, J. (2009): A Phylogenetic Interpretation of the Comparative Method. Journal of Language Relationship 2. 115-138.
  • Fox, A. (1995): Linguistic reconstruction. An introduction to theory and method. Oxford University Press: Oxford.
  • François, A. (2014): Trees, waves and linkages: models of language diversification. In: Bowern, C. and B. Evans (eds.): The Routledge handbook of historical linguistics. Routledge: 161-189.
  • Georg, S. (2017): The Role of Paradigmatic Morphology in Historical, Areal and Genealogical Linguistics. Journal of Language Contact 10. 353-381.
  • Glück, H. (2000): Metzler-Lexikon Sprache . Metzler: Stuttgart.
  • Gray, R. and Q. Atkinson (2003): Language-tree divergence times support the Anatolian theory of Indo-European origin. Nature 426.6965. 435-439.
  • Harrison, S. (2003): On the limits of the comparative method. In: Joseph, B. and R. Janda (eds.): The handbook of historical linguistics. Blackwell: Malden and Oxford and Melbourne and Berlin. 213-243.
  • Haspelmath, M. and U. Tadmor (2009): The Loanword Typology project and the World Loanword Database. In: Haspelmath, M. and U. Tadmor (eds.): Loanwords in the world’s languages. de Gruyter: Berlin and New York. 1-34.
  • Hennig, W. (1950): Grundzüge einer Theorie der phylogenetischen Systematik. Deutscher Zentralverlag: Berlin.
  • Hoenigswald, H. (1960): Phonetic similarity in internal reconstruction. Language 36.2. 191-192.
  • Hoijer, H. (1956): Lexicostatistics. A critique. Language 32.1. 49-60.
  • Jarceva, V. (1990): . Sovetskaja Enciklopedija: Moscow.
  • Klimov, G. (1990): Osnovy lingvističeskoj komparativistiki [Foundations of comparative linguistics]. Nauka: Moscow.
  • Lehmann, W. (1969): Einführung in die historische Linguistik. Carl Winter:
  • List, J.-M. (2016): Beyond cognacy: Historical relations between words and their implication for phylogenetic reconstruction. Journal of Language Evolution 1.2. 119-136.
  • Makaev, E. (1977): Obščaja teorija sravnitel’nogo jazykoznanija [Common theory of comparative linguistics]. Nauka: Moscow.
  • Matthews, P. (1997): Oxford concise dictionary of linguistics . Oxford University Press: Oxford.
  • Rankin, R. (2003): The comparative method. In: Joseph, B. and R. Janda (eds.): The handbook of historical linguistics. Blackwell: Malden and Oxford and Melbourne and Berlin.
  • Sankoff, D. (1969): Historical linguistics as stochastic process . . McGill University: Montreal.
  • Weiss, M. (2014): The comparative method. In: Bowern, C. and N. Evans (eds.): The Routledge Handbook of Historical Linguistics. Routledge: New York. 127-145.