Showing posts with label Cophylogeny. Show all posts
Showing posts with label Cophylogeny. Show all posts

Monday, August 4, 2014

A network of cheese rind microorganisms?


Cheese making is about 8,000 years old, and there are now about 1,000 distinct types of cheese throughout the world. As with most ancient crafts, the art of making cheese is to get the microbes to do most of the work for you.

To this end, there has been much interest in the microbial communities that occur in cheese rinds (the bit around the outside). Different communities are expected to be associated with different styles of cheese, since the production process can be quite different. This is shown in the first figure, which emphasizes that much of the difference between cheeses is due to different maturation procedures.

From Wolfe et al. (2014).

Recently, Wolfe BE, Button JE, Santarelli M, and Dutton R (2014. Cheese rind communities provide tractable systems for in situ and in vitro studies of microbial diversity. Cell 158: 422-433) had a look at the dominant genera of bacteria and microfungi in the rind communities of 137 different types of cheese. They don't actually tell us much about which cheeses these were, merely claiming:
We attempted to evenly sample across rind type (24 bloomy rind cheeses, 52 washed rind cheeses, and 61 natural rind cheeses) and geographic regions (87 European cheeses across 9 countries; 50 American cheeses across 13 states from the West Coast to the east Coast). We also attempted to sample across different milk types (77 cow milk, 34 goat milk, 21 sheep milk, and 5 mixed milk) and milk treatments (99 raw milk, 38 pasteurized).
Based on sequencing the bacterial 16S and fungal ITS loci, the authors identified 14 bacterial and 10 fungal genera (moulds and yeasts) that occurred with an average abundance of >1%, as shown in the next figure.

The 137 rind samples with their bacterial (middle row) and fungal (bottom row) genera indicated
by different colours. The order of the samples was determined by UPGMA clustering (top row).

The authors also used shotgun metagenomic sequencing to identify a range of genes in the microorganisms. They present a phylogeny of one particular gene (shown in the next figure) that shows a close relationship between some of the cheese microbes and marine bacteria:
The widespread distribution and high abundance of marine-associated gamma-Proteobacteria, enriched in both washed and bloomy rind cheeses, was an unexpected finding in our survey of taxonomic diversity ... One possible source of these marine microbes is the sea salt used in cheese production.
[Note: the other cheese rind bacterium shown in the phylogeny, Brevibacterium linens, is the one responsible for the unbelievable smell of washed-rind cheeses such as Epoisses, Münster and Limburger. It is also responsible for personal-hygiene issues such as foot odour. You can imagine how it first got into cheese making!]


However, Ropars J, Cruaud C, Lacoste S, and Dupont J (2012. A taxonomic and ecological overview of cheese fungi. International Journal of Food Microbiology 155: 199-210), in a related study, have pointed out the usual problem with microbial phylogenies: gene trees are frequently incongruent. So, the gene phylogeny shown above is not likely to be the species phylogeny. It would thus be of great interest to investigate the full microbial network, rather than looking at a single tree.

Wednesday, June 12, 2013

Cophylogenetic networks

I want to talk about relationships between phylogenetic networks, with respect to the cophylogeny problem.  This is a problem in which phylogenetic histories, which are typically trees, share some ecological link, which may be quite strict or relatively weak, such that their evolutionary dynamics are not independent.  Given that nothing evolves in isolation, that there are more parasites than non-parasites, and that there are many genes – each of which can have its own phylogenetic history – for each organism that houses them, this is a ubiquitous scenario and one worthy of analysis.

A simple case serves to illustrate here:

The tree on the left is labelled H for the "host" phylogeny but it could equally well be named the "species" phylogeny; the tree on the right is labelled P for the "parasite" or "pathogen" phylogeny but could also be thought of as the "gene" phylogeny.  Reconciling these two, given the observed associations at the tips, is a computationally horrible problem (Libeskind-Hadas & Charleston 2010, Ovadia et al. 2011) already.  The proof Ran Libeskind-and I did (he mostly did!) first allowed the host phylogeny to be a network but the later proof by Ovadia et al. didn't.  But it's an interesting problem, because there are hybridization events of host species though not so much of individual genes.

The cophylogeny problem is best attacked as a mapping problem (yes, this is my opinion: there are others).  Given P, H and leaf associations phi, what is the minimal cost mapping of P into H that preserves phi and the structure of P and H, and which is interpretable in a well-defined way? (see Cophylogeny blog for more details.)

We typically have four event types:


codivergence : a generalisation of cospeciation, where a node in P bifurcates at the same moment as does its host node in H;

duplication : where the parasite bifurcates without a corresponding host bifurcation, such as for a gene duplication;

host switch : a parasite/pathogen establishes on a new host lineage); and
loss : we fail to see a parasite/pathogen where we expected it, caused by "missing the boat" / lineage sorting...

extinction  (above) or sampling failure.
We could in principle extend this to deal with failure-to-codiverge events:


... but these cause new headaches.


Mike Steel once very usefully asked me, what are the desired properties of cophylogeny maps?  In trying my best to answer him I realised that none of the properties really needs either phylogeny to be a tree. Nodes in P are mapped to nodes or edges in H, and the evolutionary history they imply is based on the route through H from parent node p' to child node p in P.  If H is a tree then this is unique; otherwise there can be ambiguity and a potential explosion in number of solutions.  But it does mean that potentially we can solve the mapping problem moderately well so long as P and H are at least DAGs.  While this is possible in principle, in practice, it's pretty much impossible.
It's hard enough just with trees:

Figure from Ramsden et al. 2009 Supp. material

The figure above was cut from the manuscript and relegated to supplementary material because it was kind of unnecessary, but it's very pretty so it should see the light of day.  It shows a consensus of 15 solutions / maps we found that could best explain the relationships between hosts and circoviruses; thicker lines for more frequently occurring components and different colours for different groups of maps.

The cophylogeny problem is a fun one that presents lots of modelling, computational and representational challenges.