Showing posts with label Errors. Show all posts
Showing posts with label Errors. Show all posts

Monday, February 1, 2016

Tardigrades and phylogenetic networks


In this blog we have always championed the use of Exploratory Data Analysis prior to phylogenetic analyses. This approach explores the characteristics of the data before making formal inferences about possible evolutionary scenarios. One of the reasons for doing this is the possibility of data errors. That is, we need to distinguish between estimation errors deriving from our experimental procedures and real biological scenarios, because both of these will result in complex patterns in our data.

One possible classification of the potential causes of complex data patterns in phylogenetics is this:

Estimation errors
(i) incorrect data
— inadequate data-collection protocol
— poor laboratory / museum / herbarium technique
— lack of quality control after data collection
— misadventure
(ii) inappropriate sampling
— distant outgroup
— rapid evolutionary rates
— short internal branches
(iii) model mis-specification
— wrong assessment of primary homology
— wrong substitution model
— different optimality criteria

Biological complexity
(iv) analogy
— parallelism
— convergence
— reversal
(v) homology
— deep coalescence
— duplication–loss
— hybridization
— introgression
— recombination
— horizontal gene transfer
— genome fusion

The scientific literature has a number of prime examples where people have asserted a case of biological complexity that has subsequently been questioned, and attributed to estimation errors instead.

For example, many of you will have noted the recent attention given to the release of various genome sequences from the Tardigrades, a group of microscopic animals often alleged to be the world's most resistant to environmental conditions. Two rival papers have appeared:
Thomas C. Boothby et al. (2015) Evidence for extensive horizontal gene transfer from the draft genome of a tardigrade. Proceedings of the National Academy of Sciences of the USA 112: 15976–15981.
Georgios Koutsovoulos et al. (2015) The genome of the tardigrade Hypsibius dujardini. BioRxiv preprint 33464. [Now published as: Georgios Koutsovoulos et al. (2016) No evidence for extensive horizontal gene transfer in the genome of the tardigrade Hypsibius dujardini. Proceedings of the National Academy of Sciences of the USA]
The former paper attributes their observed phylogenetic complexity to horizontal gene transfer (group v in the list above) while the latter attributes it to sequencing errors (group i). This situation is discussed in more detail elsewhere on the web, for example:
Rival scientists cast doubt upon recent discovery about invincible animals
How did these indestructible pond critters get their genes?
This difference in possible cause (of complexity) matters particularly for the use of phylogenetic networks, because both estimation errors and biological complexity will appear as reticulation patterns in any network. This is particularly important for the assertion of evolutionary scenarios such as horizontal gene transfer, because usually the only evidence for any such gene flow is the complexity of the phylogenetic network — that is, there is no independent experimental evidence, and we are relying entirely on the phylogenetic pattern analysis. Estimation errors must thus be eliminated prior to the phylogenetic analysis, if we are to produce a high quality network.

The current situation potentially has unfortunate consequences. For example, there are continual comments that horizontal gene flow is rare, particularly from zoologists, even though there is a large amount of evidence to the contrary. Situations like the current one can only add fuel to this argument, if strong claims of gene flow turn out to be erroneous. There is no quantitative basis for an assertion that gene flow is rare in zoology — those who have looked for reticulate evolution in animals have found it, and those who haven't haven't.

In the end, data-display networks are useful for displaying incongruent data patterns, but the source of the incongruence needs to be identified before these networks are turned into evolutionary networks (either explicitly drawn or verbally implied).

Monday, November 24, 2014

When infographics go wrong


Infographics have become very popular in recent decades, with the advent of computer graphics packages. Infographics combine data and pictures, trying to produce an aesthetically pleasing but still informative presentation of numeric information. Recently, the following book appeared:
The Infographic History of the World (2013)
by Valentina D’Efilippo & James Ball
HarperCollins (UK) / Firefly Books (US)


A selection of the the infographics can be perused at the senior author's web page:
http://www.valentinadefilippo.co.uk/ihw/
At the visual.ly blog the author also explains her intentions:
The Infographic History of the World is a new book that continues to push the field of infographics forward.
Our task required research, organization and the selection of topics. Then, we needed to decide how to display data in order to tell a coherent and compelling story. We have never considered this to be an alternative to tons of books of history, but hopefully a refreshing interpretation of what history is about.
With this book, we hope to lead readers on a journey, to interpret the data and find the implications that resonate with them. We don’t pretend that every set of data presents an unquestionable truth. And, rather than looking to define the world’s history, we were looking to present readers with an unconventional interpretation of the subject.
Sadly, these good intentions have not always been achieved. As noted by a review at Amazon:
the book showcases *clever* ways of displaying data, not *clear* ways of displaying it ... Far too often I had to pore over the graphic to figure out what it was trying to say.
What is worse for the readers of this blog, the information is not always correct. Consider this version of the Tree of Life, which has a long-standing tradition in systematics as one of the world's first examples of an infographic:

Click to enlarge.

Quite a number of the taxonomic labels are misplaced. You can check them for yourselves, but here is a selection of some of the surprising information contained in this infographic:
  • Amphibians are not Tetrapods
  • Humans are not Mammals
  • Mammals are not Amniotes
  • Turtles are not Reptiles
  • Lobe-finned fishes are not Sarcopterygians
  • Ray-finned fishes are not Bony Vertebrates
  • Charophytes are Land Plants
  • Hornworts are Vascular Plants
  • Ferns and Horsetails are not only Seed Plants they are Gymnosperms
  • Conifers, Gnetophytes, Gingko and Cycads are not Gymnosperms
 Clearly, little has been done to check the veracity of the information in this infographic, which completely defeats its purpose.

Wednesday, September 18, 2013

Checking data errors with phylogenetic networks


Data-display networks can be used for a number of purposes, for example: Exploratory data analysis, Displaying data patterns, Displaying data conflicts, Summarizing analysis results, and Testing phylogenetic hypotheses. One of the more important, but currently under-valued, purposes is detecting data errors.

For instance, networks can help you detect data-sampling errors or outliers (eg. wrong specimen identification, diseased specimens), as well as data-collection errors (eg. extracting the wrong DNA, amplifying the wrong gene, sequencing artifacts) and data-processing errors (eg. data entry mistakes, incorrect alignment). These types of errors will likely show up as reticulations in a network, especially a splits graph.

Perhaps the most powerful use of such networks is in conjunction with a database of gold-standard or benchmark sequences. Comparison of all new sequences with the database would allow for a systematic quality check, because the network structure of the database is already known, and any deviation from this structure highlights potential problems ("identifying idiosyncrasies that cannot be attributed to natural evolutionary processes") or indicates novel sequence variation. Much of this process can be effectively automated by computer scripts.

To date, the champion of this use of networks has been Hans-Jürgen Bandelt, who has presented a number of interesting practical examples over the past dozen years. Below, I have included an annotated list of some of the more interesting publications in this area.

Bandelt H-J, Lahermo P, Richards M, Macaulay V (2001) Detecting errors in mtDNA data by phylogenetic analysis. International Journal of Legal Medicine 115: 64-69. —The first to suggest phylogenetic analysis as a component of data-quality checking, although networks are not explicitly mentioned

Bandelt H-J, Quintana-Murci L, Salas A, Macaulay V (2002) The fingerprint of phantom mutations in mitochondrial DNA data. American Journal of Human Genetics 71: 1150-1160. — The first to explicitly suggest using networks, and then use median and quasi-median networks to detect errors in published human mtDNA control-region datasets

Bandelt HJ, Kivisild T (2006) Quality assessment of DNA sequence data: autopsy of a mis-sequenced mtDNA population sample. Annals of Human Genetics 70: 314- 326. — Use quasi-median networks to detect errors in a published human mtDNA control-region dataset

Bandelt HJ, Dür A (2007) Translating DNA data tables into quasi-median networks for parsimony analysis and error detection. Molecular Phylogenetics and Evolution 42: 256-271. — Discuss the use of quasi-median networks for error detection, and re-visit the analysis of Bandelt and Kivisild (2006)

Parson W, Dür A (2007) EMPOP — A forensic mtDNA database. Forensic Science International: Genetics 1: 88-92. — Use quasi-median networks to detect mtDNA errors in forensic data by comparison with a benchmark database

Kong Q-P, Salas A, Sun C, Fuku N, Tanaka M, Zhong L, Wang C-Y, Yao Y-G, Bandelt H- J (2008) Distilling artificial recombinants from large sets of complete mtDNA genomes. PLOS One 3: e3016. — Use median networks to detect possible artificial recombinant sequences in molecular databases (ie. chimeric sequences resulting from laboratory-induced errors)

Bandelt H-J, Yao Y-G, Bravi CM, Salas A, Kivisild T (2009) Median network analysis of defectively sequenced entire mitochondrial genomes from early and contemporary disease studies. Journal of Human Genetics 54: 174-181. — Use median networks to detect possible errors in human mtDNA genomes intended to find sequence mutations associated with particular diseases