Showing posts with label inference. Show all posts
Showing posts with label inference. Show all posts

Wednesday, June 8, 2016

Why do so few biologists look at their phylogenetic data?


Most data analyses involve processing the data using some model. For example, standard parametric statistical tests assume a normal distribution for the "error" term, as well as equal variances and linear relationships between the variables. If these model assumptions ado not hold, then any inferences from the tests may be incorrect.

It is possible to look at any dataset in a model-free manner, although this does not necessarily lead to any strong inferences. Looking at data is usually called exploratory data analysis. This is often done using graphs of various types.

Exactly the same principle applies to phylogenetics. A phylogenetic tree is an inference from the data via a given model. The inference is a reconstructed genealogical history assuming a divergent tree. In this context, different models will often (usually?) give different inferences.

Therefore, most phylogeneticists never actually see their data. What they see, instead, is the data as processed through some model. That is, they see inferences from the model, not the original data. Models are important, but the data should be even more important, for a scientist.

It is thus interesting that so many phylogeneticists skip the step of looking at their data, and proceed immediately to the model-based inference. So many of the disagreements throughout the literature end up being about the models and not the data. There are very strong opinions about which models should be used, with less attention being paid to whether the data contain sufficient information to answer the original scientific question in the first place.

A specific example of this was discussed in some earlier blog posts:
Conflicting placental roots: network or tree?
Why are there conflicting placental roots?
In this example there are three possible genealogical patterns, each of which has been reported to receive strong support from model-based tree inference of nucleotide sequences. However, when looking at the sequence data themselves, in a model-free manner using data-display networks, any one dataset shows all three possible patterns. So, any inference of a single tree is coming from the model not from the data. That is, the data do not distinguish between the three genealogies, but the models do discriminate amongst them.

It is worth mentioning here that a haplotype network is not a genealogy. Instead, it is a summary of a population dataset, which may contain some phylogenetic patterns or it may not. So, a haplotype network is closer to exploratory data analysis than it is to model-based inference. This point is clearly made by Jessica W. Leigh and David Bryant (2015. PopART: full-feature software for haplotype network construction. Methods in Ecology and Evolution 6: 1110-1116):
The haplotype networks do provide, however, a concise and accessible representation of the data themselves, one aspect which is often lost in methods heavily dependent on model-based inference.
Looking at the data before you start processing it can be a very good idea. After all, you may be able to avoid unlikely inferences.

Monday, February 16, 2015

An Hennigian analysis of the Eukaryotae


As usual at the beginning of the week, this blog presents something in a lighter vein.

Homologies lie at the heart of phylogenetic analysis. They express the historical relationships among the characters, rather than the historical relationships of the taxa. As such, homology assessment is the first step of a phylogenetic analysis, while building a tree or network is the second step.

With a colleague (Mike Crisp, now retired), I once wrote a tongue-in-cheek article about how to mis-interpret homologies, and the consequences of this for any subsequent tree-building analysis. This article appeared in 1989 in the Australian Systematic Botany Society Newsletter 60: 24–26. Since this issue of the Newsletter is not online, presumably no-one has read this article since then. However, you should read it, and so I have linked to a PDF copy [1.2 MB] of the paper:
An Hennigian analysis of the Eukaryotae


Monday, August 25, 2014

The evolution of statistical phylogenetics



For those of you who do not understand the notation:
Homo apriorius ponders the probability of a specified hypothesis, while Homo pragamiticus is interested by the probability of observing particular data. Homo frequentistus wishes to estimate the probability of observing the data given the specified hypothesis, whereas Homo sapiens is interested in the joint probability of both the data and the hypothesis. Homo bayesianis estimates the probability of the specified hypothesis given the observed data.

Wednesday, May 29, 2013

Should phylogenetic modelling proceed from simple to complex or vice versa?


In statistical model testing, models can be tested by starting with the simplest model and progressively adding model complexity until the desired level of model fit is achieved. Alternatively, one can start with the most complex model and progressively delete unnecessary components while maintaining the desired level of model fit. The first approach is constructive, in the sense that the model is constructed piece by piece (stepwise addition), while the second approach is reductive, in the sense that the full model is pared down to its simplest form (stepwise deletion).

This distinction in approaches to modelling is relevant to the difference between using trees and networks as phylogenetic models.

At the moment, the most common approach to phylogenetic analysis is the constructive one. One starts with the simplest model, a bifurcating tree, and assesses the degree to which it fits the data. If the fit is poor, as it often is with multi-gene data, especially if the gene data are concatenated, then complexity is added. For example, one might include incomplete lineage sorting (ILS) in the model, which allows the different genes to fit different trees, while still maintaing the need for a single dichotomous species tree. Alternatively, one might consider gene duplication-loss as a possible addition to the model, which is another major source of incompatibility between multi-gene data and a single species tree. Only if these additional complexities also fail to attain the desired degree of fit does one consider adding components of reticulate evolution to the model, such as hybridization or horizontal gene transfer (HGT).

The reductive (or simplification) approach, however, proceeds the other way. A general network model is used as the starting point. The various components of this model would include a dichotomous tree as a special case, along with ILS, duplication-loss, hybridization, and HGT as individual components. These special cases are evaluated simultaneously, and each one is dropped if it is contributing nothing worthwhile to the model fit. The final model consists of the simplest combination of components that still maintains the specified fit of data and model; this may indeed be a simple tree.

The main advantage of the latter approach is that all of the components of the model are evaluated simultaneously, so that their potential interactions can be quantitatively assessed. Components are dropped from the model only if they contribute nothing to the model, either independently or in synergy with the other components. That is, they are dropped only if they can be shown to be redundant.

This does not happen with the constructive approach to modelling. Here, the components are evaluated in some specified order, and components that are later in the order will not be evaluated unless the earlier components prove to be inadequate. These later components are thus potentially excluded from statistical consideration. This means that their possible contribution to biological explanation may never be quantitatively assessed.

So, in practice, evolutionary reticulation is considered to be a "last resort" in current phylogenetic analyses. It is considered as a possible biological explanation only if all else has already failed.

This philosophy seems to be as much a historical artifact as anything else. The first phylogenetic diagrams (by Buffon and Duchesne) were networks not trees, but they were replaced a century later by the tree model suggested by Darwin; and the tree has retained its primacy since that time. This leads naturally to the constructive approach to modelling, which is so prevalent in the current literature.

However, there is no necessary statistical superiority of the constructive approach to modelling. Indeed, statisticians seem to consider forward and backward selection of model components to be essentially equivalent, although they may lead to different models for any given dataset. The most commonly specified advantage of the constructive approach to modelling is that it is likely to avoid possible problems arising from having too many components in the model.

Nevertheless, the reductive approach has the distinct advantage of simultaneously evaluating all possible special cases of a network, and thus does not exclude any possible biological explanation that might apply to the observed data. This may provide more biological insight than does the construcive approach to phylogenetic modelling.

Monday, April 22, 2013

Personal Type I error rates


As usual at the beginning of the week, this blog presents something in a lighter vein. However, this week we depart from phylogenetic networks in the strict sense, and take a humorous look at the broader statistical life of biologists.

Statistics is a curious thing, which allows scientists to make probability errors of two types: Type I (also known as false positives) and Type II (also known as false negatives). Importantly, these errors can accumulate in any one experiment, so that we can also recognize an Experimentwise Error Rate, which is the sum of the individual errors associated with each experimental hypothesis test.

However, what is not widely recognized is that these errors apply in life, as well. In particular, biologists accumulate statistical errors throughout their lives, so that we all have a Personal Lifetime Error Rate.

I once wrote a tongue-in-cheek article about the accumulation of Type I errors throughout the working life of a biological scientist, and the consequences for the experiments conducted by that scientist. This article appeared in 1991 in the Bulletin of the Ecological Society of Australia 21(3): 49–53, which means that I used an ecologist as my specific example of a biologist. The principle applies to all biologists, however.

Since this issue of the Bulletin is not online, presumably no-one has read this article since 1991, although it has recently been referenced on the web (see the sixth comment on this blog post).** You, too, should read it, and so I have linked to a PDF copy [1.7 MB] of the paper:
Personal Type I error rates in the ecological sciences


** Note that I am alternately referred to as an "inveterate mischief maker" and "a very wise man"!

Monday, December 10, 2012

Data enrichment in phylogenetics


Since this is post #100 in this blog, I thought that we might celebrate with something humorous. Since evolutionists often have a tough time, this post is about how to get more out of your phylogenetic analyses than you previously thought was possible.

In 1957, Henry R. Lewis published an article about The Data-Enrichment Method (Operations Research 5: 551-554). This method was intended "to improve the quality of inferences drawn from a set of experimentally obtained data ... without recourse to the expense and trouble of increasing the size of the sample data." This distinguishes the method from similarly named techniques, such as the likelihood method of Data Augmentation, which require actual data.

Clearly, such a method is of great interest to all empirical scientists, especially those without much grant money. Indeed, The Data Enrichment Method was immediately expanded by other interested parties (see Operations Research 5: 858-859, and 6: 136), who pointed out that it can be applied iteratively to great effect, and that it can be used to support an hypothesis and also its opposite.

The important requirements for the Data Enrichment Method are: (i) a nested set of data patterns, and (ii) an a priori expectation about what should be the answer to the experimental question. All scientists should have the latter, of course, since they are supposed to be testing the expectation by calling it an "hypothesis".

Most interestingly for us, phylogeneticists will often be able to meet requirement (i), as well, because their data often form a nested set, representing the shared derived character states from which a phylogenetic tree will be derived. I therefore once wrote an article examining the application of The Data Enrichment Method to phylogenetics, where it does indeed work very well. You do need at least some data to start with, and so it does not free you entirely from the inconvenience and embarrassment of uncontrollable empirical results.

This article appeared in 1992 in the Australian Systematic Botany Society Newsletter 71: 2–5. Since this issue of the Newsletter is not online, presumably no-one has read this article since then. However, you should read it, and so I have linked to a PDF copy of the paper:
A new method for increasing the robustness of cladistic analyses

After reading it, you might like to think about how to apply this method to phylogenetic networks. The mixture of horizontal gene flow with vertical descent breaks the simple nested data pattern of a phylogenetic tree, which complicates the application of data enrichment to networks.

Wednesday, November 7, 2012

Explanation of the many names for types of phylogenetic networks


Two types of phylogenetic network are commonly recognized, although there can be gradations between the two extremes. These go by many different names, which inevitably leads to some confusion on the part of users.

Some of the names are listed here, along with an explanation of what the terminology is intended to convey. The terms are arranged in pairs, indicating the two different types of network. The "network" part of the name is assumed in each case unless indicated otherwise.

      Type 1       Type 2
  1. Affinity  Genealogical
  2. Data-display Reticulogeny
  3. Implicit  Explicit
  4. Directed  Undirected
  5. Rooted  Unrooted
  6. Splits graph Augmented tree, Reconciliation, Recombination,
                  Hybridization
1.  This reflects the biologists' perspective, describing the different purposes for which networks have been used. Affinity networks display overall similarity relationships among the organisms, whereas genealogical networks display only historical relationships of ancestry.

2.  This reflects the assumptions used for the data analysis. Data-display networks are interpreted solely as visualizations of the patterns of variation in the data, while the reticulogenies are based on some inferences about those data patterns (such as their possible cause). Some network types, such as Reduced Median Networks and Median-Joining Networks, are based on algorithms that make partial inferences from the data. Data-display networks have mainly been used as affinity networks and reticulogenies as genealogical networks.

3.  This reflects the computational perspective, describing the goal of the algorithm used to analyze the data. Explicit networks are intended to provide a phylogeny in the traditional sense used for phylogenetic trees, displaying both vertical and horizontal patterns of descent with modification. Implicit networks provide information that can be used to explore phylogenetic patterns in a dataset without any direct interpretation as necessarily showing a phylogeny. Implicit networks have mainly been used as data-display networks and explicit networks as reticulogenies.

4.  This reflects the mathematical interpretation of networks as line graphs. In a directed graph the edges have a direction, usually indicated by an arrow, in which case the edges are more correctly referred to as arcs. Undirected graphs do not have directed edges.

5.  This reflects the tree-thinking view of phylogenetic networks, in which directed graphs are called rooted trees and undirected graphs are called unrooted trees. Rooted networks are usually treated as explicit networks and are thus used as genealogical networks, although there is no reason why they could not be used simply as a convenient form of data display.

6.  This reflects the modelling approach to network analysis based on mathematical structures. Splits graphs model phylogenetic patterns as bipartitions of the data, and build the network from those partitions (the result will be a tree if there are no incompatible bipartitions). Augmented trees are essentially trees with a few added reticulation edges / arcs, while reconciliation networks are based on reconciling the differences between trees. Recombination networks are based on analyzing data patterns in terms of a simple model of genetic cross-over, while hybridization networks model the data in terms of patterns in conflicting trees.

So, there are reasons why so many different terms have appeared in the literature. Unfortunately, they are not always used consistently with the meaning that was originally intended.

Wednesday, July 25, 2012

Are mathematical constraints biologically realistic?


Mathematicians and other computational scientists have produced their own definitions of phylogenetic networks, independently of biologists. For the evolutionary type of phylogenetic network, the definition usually looks something like this:

A phylogenetic network is a rooted, directed graph (consisting of nodes, plus edges that connect each parent node to its child nodes) such that:
(1) There is exactly one node having indegree 0, the root
        - all other nodes have indegree 1 or 2
(2) All nodes with indegree 1 have outdegree 2 or 0
        - nodes with outdegree 2 are tree nodes
        - nodes with outdegree 0 are leaves, distinctly labelled
(3) The root has outdegree 2, and
(4) Nodes with indegree 2 have outdegree 1, called reticulation nodes.

An obvious question of interest is how (or whether?) this definition connects to what biologists have in mind when they use the term "phylogenetic network". Clearly, this definition places considerable restrictions on the networks that will be inferred by any mathematical algorithm, which in turn affects their use as models for biological inference.

The first thing to note is that unrooted networks are excluded, because the graph is directed. Thus, many (if not most) of the phylogenetic networks that have appeared in the literature are excluded from the discussion. Furthermore, a tree is considered to have all internal nodes with indegree 1 and outdegree 2 (i.e. no reticulation nodes), and we know this to be biologically unrealistic, in general. (Otherwise, this blog would be redundant!)

Biologically, the other parts of the definition imply:
One node of indegree 0
- the network has no previous ancestry that is to be inferred
Nodes with outdegree 0 are labelled
- observed (contemporary) taxa occur only at the leaves
All nodes with indegree 2 have outdegree 1
- reticulation and speciation cannot occur simultaneously
No nodes with indegree >2
- reticulation events cannot involve input from more than 2 parents simultaneously
No nodes with outdegree >2
- speciation involves only two children at a time.

These do not appear to be onerous biological restrictions. Indeed, he first two have been standard characteristics of tree-building for several decades. The other three are also logical extensions of  the restrictions that have previously been placed on trees. However, phylogenetic history is unlikely to have been as simple as implied by these features. Thus, biologists will need to keep a careful eye on whether the simplifications are affecting the networks inferred for their particular group of organisms.

Other restrictions

In addition to the restrictions created by the definition, other topological restrictions have been used to make the mathematical inference algorithms computationally tractable. Thus, only certain sub-families of possible networks are considered by most of the computer programs. These include:
  • tree-child network, tree-sibling network
  • level-k network, galled tree
  • binary input trees for hybridization and HGT networks
  • binary characters for recombination networks.
These restrictions may be unrelated to each other; so, we can consider them separately.

Tree-child, Tree-sibling

In a tree-child network, every internal node has at least one child node that is a tree node
- ie. a reticulation event cannot be followed immediately by another reticulation event
In a tree-sibling network, every reticulation node has at least one sibling node that is a tree node
- ie. a parent cannot be directly involved in two separate reticulation events
Note that every tree-child network will also be a tree-sibling network, but not vice versa.

Algorithmically, these two restrictions may involve the addition of extra tree nodes to an inferred network, in order to satisfy the restrictions. Biologically, the question is whether real networks are this simple. Arenas et al. (2008) simulated data under the coalescent with recombination, and found that even at small recombination rates most of the networks produced were already more complex than a tree-sibling network. On the other hand, Arenas et al. (2010) analyzed real population-level data from the PopSet and Polymorphix databases using the TCS program, and found that >98% the resulting networks could be characterized as tree-sibling. So, there is cause for optimism, in the sense that the "optimum" networks algorithmically are not necessarily complex, at least for closely related organisms (ie. within species).

Level-k network, Galled tree

A network has level k if each tangled part of the network (ie. each biconnected component) contains at most k reticulation nodes (see this previous post). This is a generalization of the older notion of galled trees, in which reticulation cycles do not overlap (ie. do not share edges or nodes), as galled trees are level-1 networks. Level-k networks can also be seen as a generalization of networks with k reticulation nodes, although there may be a difference between a network with minimum level and one with a minimum number of reticulations.

Algorithmically, these restrictions have been used to guide the search for (or choice of) the "optimal" inferred network. Biologically, these notions do not seem to have been investigated, but basically they restrict how complex inferred reticulation histories can be. In particular, they restrict the complexity of any given subset of each network. It has been noted that optimizing k can easily lead to networks that look biologically unrealistic (Huson et al. 2011).

Binary input

The requirements for binary input trees and binary characters are restrictions that have been applied in the past, because they greatly reduce the complexity of the input to the network algorithms, but they are now being relaxed. Effectively, the restrictions are to fully dichotomous trees and SNP characters. These are not unusual restrictions in evolutionary analysis, but they are obviously unrealistic.

As I noted in an earlier blog post, non-binary data often reflect uncertainty in the input, rather than a strictly bifurcating history, and this is not taken into account in the network inference if the input is restricted to a binary state. In particular, it may be unnecessarily hard to construct a network (because not all of the data signals relate to reticulation), and the resulting networks may have far too many reticulation nodes.

Conclusion

It is still an open question about the extent to which we can use these topologically restricted families of mathematical networks as a basis for reconstructing biological histories. Clearly, much more work is needed to understand the connections between the mathematical restrictions and the requirements of biological modelling.

References

M. Arenas, M. Patricio, D. Posada, G. Valiente (2010) Characterization of phylogenetic networks with NetTest. BMC Bioinformatics 11: 268.

M. Arenas, G. Valiente, D. Posada (2008) Characterization of reticulate networks based on the coalescent with recombination. Molecular Biology and Evolution 25: 2517-2520.

D. H. Huson, R. Rupp, C. Scornavacca (2011) Phylogenetic Networks: Concepts, Algorithms and Applications. Cambridge University Press.

Wednesday, March 7, 2012

RECOMB-AB


Last week my attention was drawn to the forthcoming conference RECOMB-AB 2012 : First RECOMB Satellite Conference on Open Problems in Algorithmic Biology:
 
“RECOMB-AB brings together leading researchers in the mathematical, computational, and life sciences to discuss interesting, challenging, and well-formulated open problems in algorithmic biology.”

As someone working in the field of “algorithmic biology” (which, I guess, could be defined as the application of techniques from computer science, discrete mathematics, combinatorial optimization and operations research to computational biology problems) I was, predictably, immediately enthusiastic about the conference.  

However, what really caught my attention was the following paragraph:

“The discussion panels at RECOMB-AB will also address the worrisome proliferation of ill-formulated computational problems in bioinformatics. While some biological problems can be translated into well-formulated computational problems, others defy all attempts to bridge biology and computing. This may result in computational biology papers that lack a formulation of a computational problem they are trying to solve. While some such papers may represent valuable biological contributions (despite lacking a well-defined computational problem), others may represent computational 'pseudoscience.' RECOMB-AB will address the difficult question of how to evaluate computational papers that lack a computational problem formulation.”

Calls-for-participation rarely strike such a negative tone. However, in this case I think the conference organizers have highlighted an extremely important point. Problems arising in computational biology are inherently complex and this entails a bewildering number of parameters and degrees of freedom in the underlying models. Furthermore, it is commonplace for computational biology articles to utilize a large number of intermediate algorithms and software packages to perform auxiliary processing, and this further compounds the number of unknowns (and the inaccuracies) in the system.

All this is, to a certain extent, inevitable. However, this complexity sometimes seems to have become an end in itself. This would be harmless except for the fact that scientists subsequently attempt to draw biological conclusions from this mass of data. Rarely is the question asked: is there actually any “biological signal” left amongst all those numbers? Would we have obtained similar results if we had just fed random noise into the system?

The fact that these questions are not posed, is directly linked to the lack of a clear and explicitly articulated optimization criterion.  In other words: just what are we trying to optimize exactly? What makes one solution “better” than another? What, at the end of the day, is the question that we are trying to answer? This is exactly what RECOMB-AB is getting at with the sentence, “This may result in computational biology papers that lack a formulation of a computational problem they are trying to solve”. The articulation might be slightly formal, but the point they raise is nevertheless fundamental.

It remains to be seen what kind of a role phylogenetic networks will play at RECOMB-AB, if any. For sure, the field of phylogenetic networks continues to generate a vast number of fascinating open algorithmic problems. However, are the underlying biological models precise enough to allow us to say that we are actually producing biologically-meaningful output? Overall, I think the answer is still no. However, I think that there is reason for optimism. The field is young and evolving and it is likely that both biologists and algorithmic scientists will have a significant role in shaping its future. Hopefully this interplay will allow us to move forward on the biological front without losing sight of the need for explicit optimization criteria.