Showing posts with label EDA. Show all posts
Showing posts with label EDA. Show all posts

Monday, October 5, 2020

Rogue dinosaurs, an example from the Aetosauria


In several earlier posts (a non-comprehensive link list can be found at the end of the post), I outlined how networks, tree-sample (Consensus networks, SuperNetworks) or distance-based (Neighbor-nets) may be of practical help, especially when we study phylogenetic relationships of extinct organisms.

In this post, I will further explore this by looking at a matrix for Aetosauria (Parker 2016, PeerJ) that provides an overall (relatively) strong and unambiguous signal. [NB: The reason, I prefer to use PeerJ papers as examples is that it is one of the very few journals that is open access and has a strict open data policy — to publish there, authors have to give access to the used data.]

In the abstract of the original paper, we read the following:
Nonetheless, aetosaur phylogenetic relationships are still poorly understood, owing to an overreliance on osteoderm characters, which are often poorly constructed and suspected to be highly homoplastic. A new phylogenetic analysis of the Aetosauria, comprising 27 taxa and 83 characters, includes more than 40 new characters that focus on better sampling the cranial and endoskeletal regions, and represents the most comprehensive phylogeny of the clade to date. Parsimony analysis recovered three most parsimonious trees; the strict consensus of these trees finds an Aetosauria that is divided into two main clades: Desmatosuchia, which includes the Desmatosuchinae and the Stagonolepidinae, and Aetosaurinae, which includes the Typothoracinae.
Parker's (2016) fig. 6 shows the results of the "initial analysis" (click to enlarge, colored annotations added by me).

Systematic groups based on clades are abbreviated (see next graph for full names).

A is a "Strict component consensus" of the 30 inferred MPTs (most parsimonious trees), B the Adams consensus. C the Majority rule consensus, branch labels give percentages for branches not found in all MPTs. D a "Maximum agreement subtree after a priori pruning of one taxon (black star) within the upper clade.

Parker's (2016) fig. 7 then shows the preferred result: a "reduced strict consensus of 3 MPTs" with the red star taxon removed, and (rarely seen in dinosaur phylogeny papers) branch-support — including Bootstrap support values below 70, which are very rarely reported in the literature (from my own experience it seems that editors of systematic biology journals don't like them).


Removal of one rogue taxon (called a "wildcard" in paleozoology), Aetobarbakinoides brasiliensis, substantially reduced the number of MPTs. Nonetheless, many branches have low support, and hence also the clades (used here as synonym for monophyla) derived from them – Parker uses branch-based ("stem"-based, brackets on his tree), and node-based taxa (dots).

Low branch support may or may not matter

There are two possible reasons for low branch-support:
  • non-discriminatory signal: any alternative branching pattern receives diminishing support
  • internal signal conflict: two (or more) alternatives receive similar support.
Mapping the support on the preferred (inferred) optimal tree cannot tell us whether it's the one or the other — only Support consensus networks can visualize this. Since we are interested in the rogue, I re-ran the parsimony BS analysis (10,000 quick-and-dirty replicates, following Müller 2005, BMC Evol. Biol. 5:58) including Aetobarbakinoides brasiliensis.

Support consensus network based on 10,000 parsimony BS pseudoreplicates. Trivial splits collapsed, only splits are shown the occured in at least 20% of the BS replicates.

The decreased/low BS support within the most terminal (root-distant) subtrees, the Des'ini and Par'ini, relates to conflicting alternatives involving one or two OTUs. In the case of Des'ini, it is the affinity of Lucasuchus and NCSM 21723, while in the case of Par'ini an alternative (recognizing Tecovasuchus as sister to the remainder) is found in 1 out of three BS pseudoreplicate trees. The diminishing support for basal relationships (root-proximal branches/edges) is due to the general lack of discriminatory signal (BS any alternative < 25). However, there are very few situations in which the best-supported alternative differs much from that in the preferred tree. For instance, any alternative to a Stag'inae sister relationship has even less than BS = 24 (BS = 27 in Parker's "reduced" tree).

Our rogue, however, is not really a 'wildcard'. The scored characters simply put it much closer to the outgroup than is any other ingroup taxon. A simple explanation could be that it is a most primitive (least derived) member of the Aetosauria. Another possibility is that it lacks any critical trait needed to place it within the ingroup. Since the deep splits within the Aetosauria rely on very few character changes, we can put it in different position down here and the tree will still have the same number of inferred changes.

Trivial and non-trivial taxa

The cladograms typically shown provide limited information about the signal in the underlying matrix, its strength and weaknesses, even when not "naked" but annotated using branch-support values. Given that there are no severe overlap gaps in the data, a very quick alternative is the Neighbor-net (a necessary addition, in my opinion).

Bold edges correspond to branches (hence: clades) in Parker's preferred tree.

Using this, we can directly depict which groups, potential clades, draw substantial (partly trivial) character support.

For instance, according to Parker's tree and following cladistic classification, Stagonolepis is an invalid taxon: one species (St. robertsoni) is part of the Stag'inae clade, the other (St. olenki) is of the Des'inae clade. Character support is, however, nearly non-existent (Bremer value = 1 and BS = 7 in the original analysis; BS ≤ 20 for any competing alternative in our re-analysis). The distance network shows us why — indeed, both species are closest to each other; but, while St. robertsoni shares a critical Stag'inae character suite and, consequently, shows the highest similarity to Polesinuchus, St. olenki does not share this (note the lack of a corresponding neighborhood). Furthermore, any alternative placement fits even less. Parker's tree only resolved it at sister to all other Des'inae because it didn't fit into any of the well-supported, terminal clades (prominent edge-bundles).

We can also see where we may have to deal with internal signal conflict, and how this may affect the tree inference and lead to ambiguous branch support. Take, for instance, the NCSM 21723 individual (= Gorgetosuchus pekinensis). It's clearly a Des'inae. The reason, we have ambiguous branch support for this staircase-like subtree is that NCSM 21723 is substantially more similar to the distant, equally evolved sister lineage, the Par'ini (purple edge bundle). Hence, it must be placed as sister to all other Des'inae, although it appears to represent a more derived form than Longosuchus, representing the next step towards the most-derived crown-taxon Desmatosuchus. Tecovachus is the source of topological conflict within the Par'ini because it is the least-derived taxon. Its primitiveness will be expressed by placing it as sister to all other Par'ini, while few shared, non-exclusive apomorphies are behind its position in the preferred tree (Bremer value = 1, BS = 48 in Parker's fig. 7).

While it is obvious that the matrix has no clear tree-like signal for resolving any OTU that is not part of the terminal Des'ini and Typ'inae lineages, our 'wildcard' (Aetobarkinoides) is particularly close to the outgroup while showing no affinity to anything else. If it is part of the ingroup, it represents the ancestral form, ie. shows a character suite that is primitive (derived traits may be missing because they are simply not preserved: see description of the taxon in Parker 2016). This is the reason why it acted rogue-ish in tree inferences even though it's favored phylogenetic position is clear.

Data

Parker's original matrix can be found in the supplement to the paper. An annotated ready-to-use NEXUS-formatted version (including my standard codelines for parsimony and distance bootstrapping) and the inference results used here can be found in this figshare submission, which I generated for a technical Q&A.



Here is the promised list of previous posts dealing with fossils and networks.

Monday, April 13, 2020

Do people admire your wine brand? A network analysis


Each year, the April edition of Drinks International magazine contains a supplement with a survey called The World’s Most Admired Wine Brands. A group of people are asked to vote for the wine brands they "most admire" based on the criteria that each brand should:
  • be of consistent and / or improving quality
  • reflect its region or country
  • be well marketed and packaged
  • respond to the needs and tastes of the target audience
  • have broad appeal among wine consumers.
The tenth list, for 2020, has just been released, although the award ceremony has been delayed because of the current pandemic. It is therefore worth looking at the past decade, to see what these lists look like.


The people polled each year are drawn from "a broad spectrum of the global wine community", which apparently includes: masters of wine, sommeliers, commercial wine buyers, wine importers and retailers, wine journalists, wine consultants and analysts, wine educators, and other wine professionals. There were only 60 people involved back in 2012, but there are now more than 200.

The people could originally vote for up to six wine brands, but apparently they are now asked for only three choices. Furthermore, they are provided with a list of previous winners, including "a list of more than 80 well-known brands and producers, but as usual we also encourage the option of free choices".

I have compiled the poll results for the years 2011-2020 inclusive. Each of the published lists contains only the results for the top 50 ranked wine brands in that year — all we know about the other brands is that were ranked lower than 50th place in that year. We also do not know how many people actually voted for each of the brands that did make it into the top 50.

Across the 10 years, 116 different brands have appeared at least once in the lists. However, only 9 of these brands appeared in all 10 lists, with a further 15 brands appearing in 9 of the 10 lists. There have been 36 brands (31%) that appeared only once each. There is thus a great deal of variability in "admiration" from year to year.


As usual in this blog, we can get a picture of this variability by using a phylogenetic network, as a form of exploratory data analysis. For the first analysis, I calculated the similarity of the 10 years using the Bray-Curtis distance, based on all 116 wine brands. A Neighbor-net analysis was then used to display the between-year similarities, as shown in the graph above. Years that are closely connected in the network are similar to each other based on the ranking of the wine brands, and those that are further apart are progressively more different from each other.

This graph shows a basic gradient from 2011, at the top-left, anti-clockwise around to 2020, at the top-right. So, the rankings changed progressively through time, which is not unexpected. However, the first three years, clustered at the left, are quite different from the seven later years, at the right. Indeed, one brand (Black Tower) appeared only in the first three years, while five others appeared twice there only.

Also, this year, 2020, is notably different from previous years (as indicated by the long terminal network edge). Indeed, quite a few long-standing wine brands disappeared from the list this year, including six that had appeared in every previous list. These were replaced by 15 new brands, which had never appeared before, including the brand ranked first (Catena, from Argentina).

We can look at the brands (instead of the years) by doing the same form of network analysis. To simplify things, I included only those 55 wine brands that appeared in at least 4 of the 10 lists, as shown in the next graph. Each brand is represented by a dot in the network. Brands that are closely connected in the network are similar to each other based on their rankings across the 10 polls, and those that are further apart are progressively more different from each other.

Network of the Most Admired Wine Brands from 2011-2020

Basically, the network progresses from the most highly admired brands at the top down to the less-admired wine brands at the bottom. High admiration can be achieved either by being ranked in the lists in most years, or by achieving a high ranking in at least a few years.

Clearly, the most highly admired brand is Torres (from Spain), which is marked in red in the network. It was ranked in the top 3 in every year; and, indeed, it was first or second in each of the first nine years, dropping to third this year. Penfolds (from Australia) was ranked in the top 5 every year, while Concha y Toro (from Chile; known for their Casillero del Diablo wines) was always in the top 6. Nothing else comes even close to these three brands (eg. Vega Sicilia, also from Spain, varied from 2nd to 14th).

Those brands that appeared in all 10 lists are shown in blue in the network, while those in green appeared in 9 of the 10 years. Note that some of the latter are at the bottom of the network, indicating that they rarely ranked highly, when they did appear in the lists.

Those countries that produce the most wine dominate the lists, of course, although the two biggest producers, Italy followed by Spain, do not do the best in terms of admiration. This is shown in the table of how many of the 116 wine brands come from each country.

France
Australia
Spain
USA
Italy
New Zealand
Chile
South Africa
Portugal
Germany
Argentina
Canada
China
Hungary
Lebanon
21
16
15
13
9
9
8
8
7
3
3
1
1
1
1
For Portugal, 5 of the 7 brands are based in the Port-producing region, rather than making table wine. For France, 10 of the 21 brands are from Bordeaux. Interestingly, 4 of these were among those 6 dropped from the lists for the first time this year. Apparently, admiration for the wine chateaux of Bordeaux is waning, along with declining purchases of their produce.

Monday, February 3, 2020

A network of life expectancy and body mass index


At my advanced age, the concept of Life Expectancy (the average age at which people of my generation die) becomes of some practical importance. Perhaps more importantly, the concept of Healthy Life Expectancy rears its head, this being the average age at which one's health starts to notably deteriorate.

Both of these human attributes are related to many things, but in the modern world Obesity is one of the most important contributors to lack of health. This is frequently measured as the Body Mass Index (BMI), defined as the body mass (kilograms) divided by the square of the body height (meters). A BMI > 30 is classified as Obese, and this is definitely considered to represent lomg-term poor health.

So, let's look at some data, to see how the USA currently fares with regard to these characteristics. The US Burden of Disease Collaborators recently released some up-to-date data (The state of US health, 1990-2016. Burden of diseases, injuries, and risk factors among US states. Journal of the American Medical Association 319: 1444-1472). You can consult their Table 1 if you want to consider the major recent causes of death in the USA.

However, we will focus on the positive side, instead — how long do people live? The first graph here shows the relationship between the two Life Expectancy variables for the year 2016, with each point representing one state of the USA, plus DC. The line shown on the graph represents the national average.

Life expectabcy versus healthy life expectancy

As expected, there is a high correlation between the two variables, although there is a 6-year difference in Expectancy among the various states. The top states include Hawaii, California, Connecticut, Minnesota, New York, Massachusetts, Colorado, New Jersey and Washington; while the bottom states are Mississippi, West Virginia, Alabama, Louisiana, Oklahoma, Arkansas, Kentucky, Tennessee and South Carolina. The social and economic differences between those two groups should be clear to everyone, and this is well-known to relate to life-length.

The national average for Life Expectancy is 78.9 years, while the Healthy Life Expectancy is 67.7 years (ie. 11.2 years less). This probably doesn't surprise you — the last 11 years of your life is likely to be spent dealing with ill health. The points on the graph are scattered around the national-average line except at the lowest Expectancies — this implies a shorter period of unhealth at the end of life for those with a poor Life Expectancy. Notably, Mississippi has the lowest Life Expectancy but only the 5th lowest Healthy LE.

We can now turn to look at Body Mass Index (BMI) and how it relates to Healthy Life Expectancy. This is shown in the next graph, where the BMI data refer to the percentage of people who are obese (BMI > 30). Once again, each point refers to a single state. Clearly, as Obesity increases then Healthy LE decreases. The medical people have been telling us this for decades.

Body mass index versus healthy life expectancy

Note, however, the big difference in obesity levels between the states (15.5 percentage points) — there are nearly two-thirds more obese people in some states than in others. The states with the highest Obesity levels include West Virginia, Mississippi, Oklahoma, Iowa, Alabama, Louisiana and Arkansas, while the other extreme includes Colorado, DC, Hawaii, California, Montana, Utah, New York and Massachusetts.

Also, note that the relationship between the Obesity and Life Expectancy variables is not linear. Below 26% population obesity there is little change in average Life Expectancy, whereas above 30% obesity levels Life Expectancy declines rapidly. For every 1% increase in average Obesity the average LE is reduced by 0.3 years.

Two of the territories are labeled in the graph, as showing unusual patterns. The people of the District of Columbia are clearly not "fat cats", as often depicted, but their lives are apparently not all that healthy. On the other hand, the people of Iowa somehow manage to remain healthy for longer than average, even though they have one of the highest Obesity levels.

Finally, we can put all of this together in a single network, depicting the data patterns. As usual in this blog, one of the simplest ways to get a pictorial overview of the data is to use a phylogenetic network, as a form of exploratory data analysis. For this analysis, I first calculated the similarity of the states using the manhattan distance, based on the three variables listed above. A Neighbor-net analysis was then used to display the between-territory similarities.

The resulting network is shown in the final graph. Territories that are closely connected in the network are similar to each other based on their two Life Expectancies and BMI levels, and those that are further apart are progressively more different from each other.

Network of life expectancy and body mass index

In this case, the network displays states with decreasing Life Expectancies from top to bottom, and decreasing Obesity from left to right. It makes visually clear that those states with the shortest Life Expectancies are almost always associated with high Obesity levels (ie. they are at the bottom-left of the network).

For longer Life Expectancies, some states have high Obesity levels (top-left of the network) while some have lower levels (top-right). Iowa is shown as quite distinct from the other states (it has a long edge of its own), since it has longer LE than would be expected for its population Obesity level.

Monday, January 20, 2020

Worldwide gender differences in amount of paid versus unpaid work


A few weeks ago, I wrote about National differences in the amount of paid and unpaid work. This involved a look at the time that people spend per day on each of various different activities, averaged across each year. The data came from the time-use surveys conducted by the Organisation for Economic Co-operation and Development (OECD) for its 30 member countries. I concluded that there are many similarities among countries that share strong cultural ties, although some countries stand out as unusual within this context.


Four main categories of time use are reported in the surveys: Paid Work or Study, Unpaid Work, Personal Care, and Leisure Time; these are described in more detail in my previous post. The aggregated results for each country are available online, including data for three non-OECD countries, for comparison (China, India, South Africa).

Of particular interest is that these data are actually aggregated separately for males and females (see Balancing paid work, unpaid work and leisure). This allows us to look at the various national time-management behaviors in the light of potential differences in gender roles within those countries.

Obviously, we expect some consistent gender differences, not least because in most cultures it is the females who have traditionally been the primary care-givers in a family, and this is one of the main unpaid work activities. We can use the OECD data to look at this in a bit more detail.

Overall gender differences

First, we can look at the overall time-management differences between the two genders.

In order to get an overview of the current differences between the 33 countries (30 OECD, 3 non-OECD), I have performed this blog's usual exploratory data analysis. The available data are multivariate, since there are five measured variables for each country — total paid work, total unpaid work, total personal care time, leisure time (each measured in average number of minutes per day), plus Other (to make a total of 1,440 minutes per day). One of the simplest ways to get a pictorial overview of the data patterns is to use a phylogenetic network, as a form of exploratory data analysis. For this network analysis, I first calculated the gender differences as Male time minus Female time (for each variable separately), and then calculated the similarity of the countries using the manhattan distance. A Neighbor-net analysis was then used to display the between-country similarities.

The resulting network is shown in the first graph. Countries that are closely connected in the network are similar to each other based on their average gender difference in time management, and those countries that are further apart are progressively more different from each other.


At the bottom of the network we see those countries with the biggest gender differences, progressing up to the top with those countries with the least difference.

So, the non-European countries show the most traditional separation of gender roles, with Portugal standing out as being the only one from Europe. China is not situated with the other two Asian countries (Japan, Korea), although why it should be similar to South Africa is not clear.

Indeed, the English-speaking part of the southern hemisphere does not do well, with all three countries (South Africa, Australia, New Zealand) showing stronger gender differences than any of the other English-speaking countries (Canada, USA, UK), except for the Irish (who thus have some explaining to do).

The Scandinavia countries are at the top (Sweden, Norway, Denmark), with the smallest gender differences, which will not surprise anyone who knows these people. On the other hand, the location of France may surprise those people who have a clichéd image of the behavior of Frenchmen. France is clearly separated from the more traditional societies of the other Mediterranean countries (Spain, Greece, Italy), appearing in the network with other northern countries (Belgium, Netherlands, Germany).

Finland and Estonia have strong historical ties, and they are distinct from the other Baltic countries (Latvia and Lithuania).

Work time differences

Having thus noted that there are some strong gender differences in time-management between countries, we can now proceed to look specifically at Paid versus Unpaid work.

First, we can simply take the total amount of reported Paid + Unpaid work, and compare gender differences across the various countries. This table lists the reported differences expressed as Male time minus Female time, in average minutes per day:
Norway
New Zealand
Denmark
Netherlands
Japan
Canada
USA
Australia
Germany
Mexico
Sweden
Turkey
UK
Korea
Austria
Belgium
Ireland
Poland
Luxembourg
France
Latvia
Finland
China
South Africa
Slovenia
Hungary
Lithuania
Estonia
Spain
Greece
Italy
Portugal
India
18.5
10.0
8.8
4.0
-3.3
-3.3
-4.7
-7.3
-7.8
-10.8
-11.3
-13.2
-16.1
-16.4
-17.9
-18.7
-20.1
-24.6
-27.4
-29.3
-35.1
-39.6
-44.0
-47.5
-54.1
-61.3
-65.3
-69.8
-73.7
-74.7
-87.9
-90.8
-94.3

These time differences between males and females become very large towards the bottom of the table, where in India it amounts to 1.5 hours per day, and is >1 hour for all of the bottom 8 countries. Note that only in the first four countries (out of the 33) does the total work time for males exceed that for females. It is unclear why the reported gender difference is so large for Norwegians; but maybe some of my readers might think that this could be a useful role model for the other countries!

We can now look at the balance between paid and unpaid work for the two genders. The following graph shows the difference as Male time minus Female time (in average minutes per day) for Paid work (horizontally) and Unpaid work (vertically). The pink line indicates the balance between the two types of work (ie. a decrease in paid work is balanced by a corresponding increase in unpaid work, and vice versa).

Gender differences in amount of paid versus unpaid work

The horizontal axis makes it clear that males always do more paid work than do females, on average, in every country, and up to 4 hours more in Mexico and Turkey. The vertical axis makes it clear that females always do more unpaid work than do males, on average, in every country, and up to 5 hours more in India.

These two variables must be correlated, since most people do either the one type of work or the other. However, in most countries the gender balance is not equal, as shown in the table above (females usually do more total work than do males). Some countries come close to a balance (indicated by the pink line), including the USA.

Note that the country with the closest gender equality is the one with the best reputation in this regard: Sweden. For example, Swedish couples frequently share their workplace parental leave for new-born children, so that there is very little gender bias in who is the primary care-giver in a family. However, the gender bias still amounts to 5–7 minutes of work per day, even in Sweden.

At the other end of the scale, there are a number of countries that still abide by the traditional model of gender roles, of which five are labeled at the bottom of the graph. These cover quite a diversity of cultures, so that no generalizations can be made. However, the gender bias in India exceeds that in Mexico — the Indians report less total work time than do the Mexicans, but that time is organized in a more gender-biased manner. Once again, Portugal stands out among the European countries — the Portuguese work longer hours than do other Europeans, and that time is organized in a more gender-biased manner.

Other differences

Gender differences occur among the other survey variables, as well. As one simple example, we can consider the time reported as being spent Eating & Drinking. This graph shows the time (in minutes per day) spent by the males (horizontally) and the females (vertically) for each of the 33 countries.

Gender differences in amount of time spent eating and drinking

As you can see, there is not a big difference between the two genders, in any country. However, in most countries males do report spending more time feeding themselves than do the females (ie. the points are to the right of the pink line, which represents equal time).

The Mediterranean countries spend the most time eating and drinking, with Greece showing the biggest gender difference. The fast food preferred by Canadians and Americans clearly does not take much time to consume, in any given day, and females can apparently eat it just as fast as males.

Conclusion

The conclusion surprises no-one — all countries have clear gender differences in who does most of the unpaid work. Two Scandinavian countries stand out — Norway, because males do more total work than do females; and Sweden, where the gender balance between paid and unpaid work is smallest. Some countries still show strong gender bias, including India, Mexico, Turkey and Portugal

Monday, December 30, 2019

National differences in amount of paid versus unpaid work


Countries differ in many cultural ways. An important one of those ways concerns how time is managed. There are 24 hours in every day, and 7 days in every week, and the time that people spend on each of the different activities can be averaged across each year. When combined across the whole population, these averages usually differ between countries, and this is what we mean when we recognize national behaviors. There are, however, many similarities among countries that share strong cultural ties.

The Organisation for Economic Co-operation and Development (OECD) has collected data on this matter among its member countries, as it also has for many other cultural and economic characteristics. In each of the 30 member countries, the OECD conducts regular "time-use surveys, based on nationally representative samples of between 4,000 and 20,000 people." The aggregated results are available online, including data for three other countries, for comparison (China, India, South Africa).


Four main categories of time use are reported by the surveys:
  • Paid Work or Study, which includes paid work time, time in school or classes, travel to and from work / study, research / homework, and job search.
  • Unpaid Work, which includes child care, care for other household members, care for non household members, routine housework, shopping, volunteering, and travel related to household activities.
  • Personal Care, which includes sleeping, eating & drinking, medical services, and travel related to personal care.
  • Leisure Time, which includes sports, participating / attending events, visiting or entertaining friends, and TV or radio at home.
Of particular interest is that the data are aggregated separately for males and females. I will look at the gender data in a future blog post, while here I will look only at the pooled data for each country.

National differences

In order to look at the current differences between the 33 countries (30 OECD, 3 non-OECD), I have performed this blog's usual exploratory data analysis. The available data are multivariate, since there are five measured variables for each country — total paid work, total unpaid work, total personal care time, leisure time (each measured in average number of minutes per day), plus Other (to make a total of 1,440 minutes per day). One of the simplest ways to get a pictorial overview of the data patterns is to use a phylogenetic network, as a form of exploratory data analysis. For this network analysis, I calculated the similarity of the countries using the manhattan distance; and a Neighbor-net analysis was then used to display the between-country similarities.

The resulting network is shown in the first graph. Countries that are closely connected in the network are similar to each other based on their average time management, and those countries that are further apart are progressively more different from each other.

National differences in amount of paid versus unpaid work

The expected cultural similarities of the countries are, in most cases, reflected in the network. For example, the country that we might expect to be the most different to the other 32 is Mexico (it is the only country from Latin America), and it is also the most isolated one in the network. It is characterized by having the greatest amount of average work time per day, particularly unpaid work, and the least leisure time.

Furthermore, the three Asian countries are clustered together: Japan, Korea, and China. They have similar high amounts of paid work, but much less unpaid work than the Mexicans.

On the other hand, it is not clear why India is shown as very similar to some of the European countries, given its very different culture. However, differences do appear in the gender patterns discussed below.

In other cases, there are occasional countries that are not where we might anticipate them to be in the network, given other known historical and cultural similarities, particularly language. For example, Sweden is not near the other Nordic countries (Denmark, Norway, Finland), as the Swedes report many more minutes of paid work per day, and correspondingly less time on each of the other activities. Portugal is not near Spain, Italy, and France, as they also report more minutes of paid work, and specifically less leisure time. On the other hand, Australia is not near Canada, the USA, New Zealand, and the UK, because the Australians report fewer minutes of paid work but correspondingly more unpaid work time. The people of Latvia and Lithuania also report many more minutes of paid work per day than do those of Poland and Estonia.

Other differences

Lest you get the impression that historical and cultural ties dominate the time-management data, we can look at one part of the data in detail.

As noted above, the Personal Care data includes separate information for sleeping versus eating & drinking. In the next graph I have plotted these two variables against each other (in average minutes per day), for all 33 countries.

Time spent sleeping versus eating & sleeping for 33 countries

As you can see, thee is no correlation whatsoever between these two variables. That is, extra eating and drinking time does not come out of the time allocated for sleeping, or vice versa.

Moreover, you will note that the denizens of the three Asian countries do not behave anything like each other, particularly as the Chinese sleep longer than everyone except the South Africans. Nor do the Swedes behave much like the Danes, in terms of eating and drinking.

Finally, the Mexicans report that they do not spend much time eating or sleeping, which follows from the work data discussed above. Instead, it is the Mediterranean peoples who like to spend their time eating and drinking. On the other hand, the Americans (and Canadians) certainly behave like they live on fast food, spending less time on eating and drinking than anyone else. They do, however, like their 8.5 hours sleep per day, which most other populations think they can do without that extra half hour.

Monday, November 11, 2019

A new playground for networks and exploratory data analysis


[This is a post by Guido with some help from David]

There tend to be two types of studies of inheritance and evolution. First, there is evolution of organisms, either of the phenotype (morphology, anatomy, cell ultrastructure, etc) or genotype (chromosome, nucleotides). The latter involves direct inheritance, but it is often treated as including all molecules, although it is the nucleotides (and chromosomes) that get inherited, not amino acids, for example.

Second, there are studies of the evolution of behaviour, which has focused mainly on humans, of course, but can include all species. For humans, this includes socio-cultural phenomena, particularly language (written as well as spoken), but also including cultural advancements such as social organization, tool use, agriculture, etc., which are inherited indirectly, by learning.

However, we rarely see studies that are multi-disciplinary in the sense of combining both physical and behavioural evolution. It is therefore very interesting to note the just-published preprint by:
Fernando Racimo, Martin Sikora, Hannes Schroeder, Carles Lalueza-Fox. 2019. Beyond broad strokes: sociocultural insights from the study of ancient genomes. arXiv.
These authors provide a review about the extent to which the analysis of ancient human genomes has provided new insights into socio-cultural evolution. This provides a platform for interesting future cross-disciplinary research.

The authors comment:
In this review, we summarize recent studies showcasing these types of insights, focusing on the methods used to infer sociocultural aspects of human behaviour. This work often involves working across disciplines that have, until recently, evolved in separation. We argue that multidisciplinary dialogue is crucial for a more integrated and richer reconstruction of human history, as it can yield extraordinary insights about past societies, reproductive behaviours and even lifestyle habits that would not have been possible to obtain otherwise.
Since multi-disciplinary dialogue is a focal point here at the Genealogical World of Phylogenetic Networks. Since our blog embraces non-biological data, we have done a little brainstorming, to put forward some ideas based on Racimo et al.'s comments. The four figures contain some extra discussion, with some visual representations of the ideas.

Why it's important to correlate genetic, linguistic and socio-cultural data. The doodle shows a simple free expansion model of a founder population with three genotypes (yellow, green, blue), a shared language (L) and two major cultural innovations (white stars). Because of drift and stochastic intra-population processes (size represent the size of the actively reproducing populace) the first expansion (light gray arrows) lead to 'tribes' that show already some variation. The smaller ones close to the founder population spoke still the same language, the ones further away used variants (dialects) of L (L', still close to L, L'', more distinct). Because of bootlenecks, geographic distance and differing levels of inbreeding (the smaller a population, the farther away from the source, the more likely are changes in genotype frequency), each population has a different genotype composition. The second expansion (mid-gray arrows) mixing two sources leads to a grandchild that evolved a new language M and lost the blue genotype. Because the cultural innovations are beneficial, we find them in the entire group. In extreme cases of genetic sorting and linguistic evolution, such shared cultural innovations may be the only evidence clearly linking all these populations.

Social-cultural character matrices

Correlating different sets of data and (cross-)exploring the signal in these data can be facilitated by creating suitable character matrices. In phylogenetics, we primarily use characters that underlie (ideally) neutral evolution, such as nucleotide sequences and their transcripts, amino-acid sequences. When using matrices scoring morphological traits, we relax the requirement of neutral evolution, but we are still scoring traits that are the product of biological evolution. However, we don't need to stop there, phylo-linguistics is an active field, even though languages involve different evolutionary constraints and processes than we meet in biology. Data-wise there are nonetheless many analogies, and phylogenetic methods seem to work fine.

So, why not also score socio-cultural traits in a character matrix? For instance, we can characterize cultures and populations by basic features including: the presence of agriculture, which crops were cultivated, which animals were domesticated, which technological advances were available, whether it was a stone-age, bronze-age, iron-age culture, etc. Linguistically, we could also develop matrices of local populations, with regional accents or dialects, etc.

Creating such a matrix should, of course, be informed by available objective information. As in the case of morphological matrices or non-biological matrices in general, we should not be concerned about character independence. We don't need to infer a phylogenetic tree from these matrices, as their purpose is just to sum up all available characteristics of a socio-cultural group.

Second phase: stabilization of differentiation pattern. While the close-by tribes are still in contact with the mother population, the most distant lost contact. As consequence the gene pools of the L/L'-speaking communities will become more similar, and new innovations acquired by the founder population (black star) are readily propagated within its cultural sphere. Re-migration from the larger M-speaking tribe to the struggling L''-speakers (small population with high inbreeding levels) lead to the extinction of the blue genotype in the latter and increased 'borrowing' of M-words and concepts.

Distance calculations

Pairwise distance matrices are most versatile for comparing data across different data sets.

First, any character matrix can be quickly transformed into a distance matrix, and the right distance transformation can handle any sort of data: qualitative, categorical data as well as quantitative, continuous data.

Second, the signal in any distance matrix can be quickly visualized using Neighbor-nets. This blog has a long list of posts showing Neighbor-nets based on all sorts of sociological data that don't follow any strict pattern of evolution, and are heavily biased by socio-cultural constraints (eg. bikability, breast sizes, German politics, gun legislation, happiness, professional poker, spare-time activities). We have even included celestial bodies.

Third, distance matrices can be tested for correlation as-is, without any prior inference, using simple statistics, such as the Pearson correlation coefficient. To give just one example from our own research: in Göker and Grimm (BMC Evol. Biol. 2008), the latter was used for testing the performance of character and distance transformations for cloned ITS data covering substantial intra-genomic diversity, by correlating the resulting individual-based distances with species-level morphological data matrices. (The internal transcribed spacers are multi-copy, nuclear-encoded, non-coding gene regions; in the simplest case each individual has two sets of copies, arrays, one inherited from the father, the other from the mothers, which may differ between but also within the individual.)

In the context of Racimo et al.'s paper, one could construct a genetic, a socio-cultural, a linguistic and a geographical matrix, determine the pairwise distances between what in phylogenetics are called OTUs (the operational taxonomic units), and test how well these data (or parts of it) correlate. The OTUs would be local human groups sharing the same culture (and, if known) language.

Alternatively, one can just map the scored socio-cultural traits onto trees based on genetic data or linguistics.

A new culture with its own language (Λ), genotype (red) and innovations (ruby-red pentagon) migrates close to the settling area of the L-people. Because of raids, genotypes and innovations from the the L-people get incorporated into the the Λ-culture.

How to get the same set of OTUs

The Göker & Grimm paper mentioned above tested several options for character and distance transformations, because we faced a similar problem to what researchers will face when trying to correlate socio-cultural data with genetic profiles of our ancestors: a different set of leaves (the OTUs). We were interested in phylogenetic relationships between individuals using data representing the genetic heterogeneity within these individuals.

Genetic studies of human (ancient or modern) DNA use data based from individuals, but socio-cultural and linguistic data can only be compiled at a (much) higher level: societies, or other groups of many individuals. In addition, these groups may also span a larger time frame. Since humans love to migrate, we are even more of a genetic mess than were the ITS data that we studied.

One potential alternative is to use the host-associate analysis framework of Göker & Grimm. Instead of using the individual genetic profiles (the associate data), one sums them across a socio-cultural unit (serving as host). The simplest method is to create a consensus of the data (in Göker & Grimm, we tested strict and modal consensuses). This produces sequences with a lot of ambiguity codes — genetic diversity within the population will be presented by intra-unit sequence polymorphism (IUSP). Standard distance and parsimony implementation do not deal with ambiguities, but the Maximum likelihood, as implemented in RAxML, does to some degree. A gapstop is the recoding of ambiguities as discrete states for phylogenetic analysis (tree and network inference) as done by Potts et al. (Syst. Biol. 2014 [PDF]) for 2ISPs ('twisps'), intra-individual site polymorphism. It can't hurt to try out whether this works for IUSPs, too.

Since humans (tribes, local groups) often differ in the frequency of certain genotypes, it would be straightforward to use these frequencies directly when putting up a host matrix. Instead of, for example, nucleotides or their ambiguity codes, the matrix would have the frequency of the different haplotypes. We can't infer trees from such a matrix (we need categorical data), but we can still calculate the distance matrix and infer a Neighbor-net.

The 'phylogenetic Bray-Curtis' (distance) transformation introduced in Göker & Grimm (2008) also keeps the information about within-host diversity when determining inter-host distances (see Reticulation at its best ...)


Transformations for genetic data from smaller to larger, more-inclusive units are implemented in the software package POFAD by Joli et al. (Methods in Ecology & Evolution, 2015. Their paper also provides a comparison of different methods, including the ones tested in Göker & Grimm (2008, also implemented in the tiny executables g2cef and pbc, compiled for any platform).

The process of assimilation. The Λ-people subdued the L-culture with the consequence that all innovations are shared in their influence sphere. Having a much smaller total population size, the language of the invaders is largely lost but the new common language L* still includes some Λ-elements (in a phylogenetic tree analysis, L* would be part of the L/M clade, using networks, L* would share edges with Λ in contrast to L and M). The L''/M-speaking remote population is re-integrated. The invaders' genotype (red) becomes part of the L-people's gene pool. Re-migration (forced or not) introduces L-genotypes into the original Λ-population. Only by comparing all available data, ideally covering more than one time period, we can deduce that the M-speakers represent an early isolated subpopulation of the L-people that was not affected by the Λ-invasion. With only the genetic data at hand, one may identify the M-speakers as one source and the Λ-tribe as another source for the L*-people, and infer that all L/M and Λ-tribes share a common origin (since the yellow genotype is found in both the M- and the original Λ-population).

Conclusion

It therefore seems to us that there is enormous potential for multi-disciplinary work, that truly combine organismal and socio-cultural evolution. We have provided a few practical suggestions here about how this might be done. We encourage you all to have try some of these ideas, to see where it leads us all.

Monday, August 12, 2019

Public transit trips in the USA


Public transport, or mass transit, has long been a politically charged issue, throughout the world. However, the modern world now recognizes that it is an effective way to deal with mass movements of people in a manner that respects the use of non-renewable resources.

After all, the only way to continue with autonomous transportation is to get rid of fossil fuels. However. electric cars will not be of much use until we work out where we are going to get all of the needed extra electricity, in a manner that is environmentally friendly. There is not much point in simply moving the burning of fossil fuels from the vehicle (ie. gasoline) to a power station that also burns fossil fuels (eg. coal). There is also a limit to how many rivers there are left to dam for hydroelectric power; and nuclear reactors have gone out of fashion (fortunately). There is also, of course, the matter of how we are going to recycle the used (lithium-ion) batteries from the cars, which is apparently a tougher proposition than recycling the electric motors themselves.


So, until we sort this out, mass transit is a viable option for most conurbations. In this context, a conurbation (or a metropolitan area) is a contiguous area within which large numbers of people move regularly, especially traveling to and from their workplace each weekday. A conurbation often involves multiple cities and towns, as defined by political administrations or contiguous urban development — many people live in one urban area but work in another.

So, naturally, governments collect data on these matters. One such data collection is the U.S. Department of Transportation's National Transit Database. The data consist of "sums of annual ridership (in terms of unlinked passenger trips), as reported by transit agencies to the Federal Transit Administration." Data for three separate modes of transit are included: bus, rail, and paratransit. The data currently cover the years 2002–2018, inclusive.

To look at the data for the 42 U.S. conurbations included, for the year 2018, I have performed this blog's usual exploratory data analysis. I first calculated the transit rate per person, by dividing the annual number of trips for each of the three modes by the conurbation population size. Since these are multivariate data, one of the simplest ways to get a pictorial overview of the data patterns is to use a phylogenetic network. For this network analysis, I calculated the similarity of the conurbations using the manhattan distance. A Neighbor-net analysis was then used to display the between-area similarities.

The resulting network is shown in the graph. Conurbations that are closely connected in the network are similar to each other based on the trip rates, and those areas that are further apart are progressively more different from each other. In this case, there is a simple gradient from the busiest mass transit systems at the top of the network to the least busy at the bottom.


The network shows us that the New York – Newark transit-commuting area (which covers part of three states) is far and away the busiest in the USA. The subway system dominates this mass transit, of course, as it is justifiably world famous, although not always for the best of reasons as far as commuters are concerned

The San Francisco – Oakland area is in clear second place. Here, bus transit slightly exceeds rail transit. Then follows Washington DC and Boston, both of which also cover parts of three states. In Boston trains out-do buses 2:1, while in Washington it is closer to 1.5:1.

Nest comes a group of four conurbations: Chicago, Philadelphia, Portland and Seattle. Two of these cover part of Washington, but in quite different ways — in Seattle the buses dominate the system 5:1 but in Portland it is only 1.5:1. Chicago and Philadelphia share buses and trains pretty equally.

At the bottom of the network there are two large groups of conurbations, one of which does slightly better than the other at mass transit use. The least-used system is that of San Juan, in Puerto Rico, perhaps not unexpectedly. Of the contiguous U.S. states, Indianapolis (IN) has the least used system, followed by Memphis (TN–MS–AR).

Moving on, we could also look at changes in the total number of transit trips (irrespective of mode) during the period for which data are available: 2002–2018. A network is of little help here. So, it so simplest just to plot the data, as shown in the next graph.


For most of the metropolitan areas there is little in the way of consistent change through time. However, there are some areas that show high correlations between the number of trips and time. These are the areas that have shown the most consistent increase in the number of transit trips from 2002–2018:
  • Chicago (IL–IN)
  • Tampa – St Petersburg (FL)
  • Baltimore (MD)
  • Denver – Aurora (CO)
  • San Francisco – Oakland (CA)
  • Memphis (TN–MS–AR)
  • San Diego (CA)
  • Cleveland (OH)
  • Providence (RI–MA)
  • Orlando (FL)
  • Indianapolis (IN)
  • New York – Newark (NY–NJ–CT)
  • Portland (OR–WA)
  • Minneapolis – St Paul (MN–WI)
Sadly, there are also areas that have shown a consistent decrease in the number of transit trips through time (2002–2018):
  • Kansas City (MO–KS)
  • Columbus (OH)
  • Riverside – San Bernardino (CA)
Presumably these are the areas where the local politicians should be looking into how to address this long-term issue.

Declining transit numbers is a topic discussed around the web; for example: Transit ridership down in most American cities. This article has a graph neatly showing the change in transit numbers from 2017 to 2018. It shows marked decreases, particularly for bus trips, while the few increases almost all involved rail travel. Is this a short-term effect, or the start of a general long-term decline?

Monday, June 17, 2019

Ockham's Razor applied, but not used: can we do DNA-scaffolding with seven characters?


One of the most interesting research areas in organismal science is the cross-road between palaeontology and neontology, which puts together a picture marrying the fossil record with molecular-based phylogenies. Unfortunately, when it comes to plant (palaeo-)phylogenetics, some people adhere to outdated analysis frameworks (sometimes with little data).

How to place a fossil?

The fossil record is crucial for neontology as it can provide age constraints (minimum ages when doing node dating) and inform us about the past distribution of a lineage. This, especially in the case of plants that can't run away from unfortunate habitat changes, can be much different than today.

The main question in this context is whether a fossil represents the stem, ie. a precursor or extinct ancient sister lineage, or the crown group, ie. a modern-day taxon (primarily modern-day genus). For instance, the oldest crown fossil gives the best-possible minimum age for the stem (root) age of a modern lineage, whereas a stem fossil can give (at best) only a rough estimate for the crown age of the next-larger taxon/clade when doing the common node dating of molecular trees (note that fossilized birth-death dating can make use of both).

There are two commonly accepted criteria to identify a crown-group fossil:
  1. Apomorphy-based argues that if a fossil shows a uniquely derived character (ie. a aut- or synapomorphy sensu Hennig) or character suite diagnostic for a modern-day genus, it represents a crown-group fossil.
  2. Phylogeny-based aims to place the fossil in a phylogenetic framework, the position of the fossil in the genus- or species-level tree (most commonly done) or network (rarely done but producing much less biased or flawed results) then informs what it is.
(We will focus on members of modern-day genera, since it becomes more trickier for higher-level taxa, see eg. my posts thinking about What is an angiosperm? [part1][part2][why I pondered about it].)

There a three basic options to place a fossil using a phylogenetic tree.
  1. Putting up a morphological matrix, then inferring the tree. A classic but due to the nature of most morphological data sets leading to a partly wrong tree as we demonstrated in some posts here on the Genealogical World of Phylogenetic Networks (hence, such analysis should always be done in a network-based exploratory data analysis framework).
  2. Putting up a mixed molecular-morphological matrix, then inferring a "total evidence" tree. This includes sophisticated approaches that use the molecular data to implement weights on the morphological traits and/or consider the age of the fossils (so-called total evidence dating approaches). Works not that bad with animal-data, provided the matrix includes a lot of morphological traits reflecting aspects of the (molecular-based) phylogeny. Doesn't work too well for plants because we usually have much fewer scorable traits, most of which are evolved convergently or in parallel. Non-trivial plant fossils love to act as rogues during phylogenetic inference.
  3. Optimise the position of a fossil in a molecular-based tree, eg. using so-called "DNA scaffold approach" (usually using parsimony as optimality criterion) or the evolutionary placement algorithm implemented in RAxML (using maximum likelihood). A special form of this approach is to first map the traits on a (dated) molecular tree, and then find the position where a fossil would fit best.

Why (standard) phylogenetic tree-based approaches are tricky

Below a simple example, including three fossils of different age (and often, place) with different character suites.


Even though none of the derived traits (blue and red "1") is a synapomorphy (fide Hennig), we can assign the youngest fossil X to the lineage of genus 1A just based just based on its unique derived ('apomorphic') character suite. Its likely a crown-group fossil of clade 1, and may inform a minimum age for the most-recent common ancestor (MRCA) of the two modern-day genera of Clade 1.
Apomorphy-wise, fossils Y and Z cannot be unambiguously placed. The red trait appears to be independently obtained in both clades, and the blue trait may have been
To discern between the options, we'd be well-advised to do character mapping in a probabilistic framework which require a tree with independently defined branch-lengths.

Just by using parsimony-based DNA-scaffolding, fossil X would be confirmed as crown-group fossil and member of genus 1A (being identical and different from all others) and fossil Z would end up as a stem-group fossil. Fossil Y, however, would be placed as sister to genus 2C (again, identical to each other and different from all others). Using Y in node dating, would then lead to a much too old divergence age for the crown-group age of Clade 2. In reality, what researchers do with such a seemingly too old fossil is not to use it by the book, as MRCA of Genus 2B and 2C, but to inform the MRCA of eg. genera 2A, 2B, and 2C assuming that the fossil's age and trait set indicate the 2C morphology is primitive within the clade or Y is an extinct sister lineage and the shared derived trait a convergence (parallelism).

Four characters, three homoplastic and one invariant, are surely not enough for DNA-scaffolding, but adding more and more characters has a catch. Easy to do for the modern-day taxa, for which we also have molecular data, the preservation of fossils limits adding many more traits; any trait not preserved in the fossil is effectively useless when placing it (including not-preserved traits in total evidence approach may, nonetheless, help the analysis). Which brings us to the real-world example just published in Science:

Wilf P, Nixon KC, Gandolfo MA, Cúneo RA (2019). Eocene Fagaceae from Patagonia and Gondwanan legacy in Asian rainforests. Science 364, 972. Full-text article at Science website.

Why one should not place a fossil using DNA-scaffolding with seven characters

Wilf et al. show (another) spectacularly preserved fossil from the Eocene of Patagonia. Personally, I think that just publishing and shortly describing such a beautiful fossil should be enough to get into the leading biological journals.

But Wilf et al. wanted (needed?) more and came up with the following "phylogenetic analysis" to argue that their fossil is a crown-group Castanoideae, a representative of the modern-day firmly Southeast Asian tropical-subtropical genus Castanopsis, and evidence for a "southern route to Asia hypothesis" (via Antarctica and Australia, both well-studied but devoid so far of any Fagaceae presence; despite the fact that the modern-day climate allows cultivating them as eg. source for commercially used wood).


Wilf et al's Fig. 3 and Table 1 suggest to me that the paper was not critically reviewed by anyone familiar with the molecular genetics of Fagaceae or phylogenetic methods in general — perhaps this is not needed, since the first author is well-merited and the second author a world-leading expert of botanical palaeo-cladistics. However, parsimony-based DNA-scaffolding can be tricky, even with a larger set of characters (see eg. the post on Juglandaceae using a well-done matrix), and using seven is therefore quite bold. Notably, of the seven characters, one is parsimony-uninformative and four are variable within at least one of the included OTUs.

Side note: The tree used as a backbone is outdated and not comprehensive. Plastid and nuclear-molecular data indicate that the castanoids Lithocarpus (mostly tropical SE Asia) and Chrysolepis (temperate N. America) may be sisters. However, the morphologically quite similar Notholithocarpus is not related to either of these, but is instead a close relative of the ubiquitous oaks, genus Quercus (not included in Wilf et al.'s backbone tree), especially subgenus Quercus. Furthermore, the (today Eurasian) castanoid sisterpair Castanea (temperate)-Castanopsis (tropical-subtropical) have stronger affinities to the (today and in the past) Eurasian oaks of subgenus Cerris. The Fagaceae also include three distinct monotypic relict genera, the "trigonobalanoids" Formanodendron and Trigonobalanus, SE Asia, and Colombobalanus from Columbia, South America. Using a more up-to-date instead of a 2-decade-old molecular hypothesis would have been a fair request during review, as would compiling a new molecular matrix to infer a tree used as backbone (currently gene banks include > 238,000 nucleotide DNA accessions including complete plastomes). This would have also enabled the authors to map their traits using a probabilistic framework, which can protect to some degree against homoplastic bias but requires a backbone tree with defined branch-lengths.

There are many more problems with the paper and its conclusions, but this critique would be content- not network-related. Let's just look at the data and see why Wilf et al. would have better off not showing any phylogenetic analysis at all (and the impact-driven editors and positive-meaning reviewers should have advised against it). Or a network.

Clades with little character support

The scaffolding placed the Eocene fossil in a clade with both representatives of Castanopsis, from which it differs by 0–2 and 1–4 traits, respectively. Phylogeny-based, the fossil is a stem- or crown-Castanopsis.

However, the fossil has a character suite that differs in just a single trait (#6: valve deshiscence) from the (genetically very distant) sister taxon of all other Fagaceae, Fagus (the beech), used here as the outgroup to root the Castanoideae subtree. As far as apomorphies are concerned, the data are inconclusive as to whether the fossil represents a stem-Castanoideae (or extinct Fagaceae lineage) or a Castanopsis — this critical, potentially diagnostic derived trait, partial valve dehiscence, is only shared by the fossil and some but not all modern-day Castanopsis. This particular trait is not mentioned elsewhere in the text, although it is the reason the fossil is placed next to Castanopsis and not the outgroup Fagus in the "phylogenetic analysis".

In the following figure, I have mapped (with parsimony) the putative character mutations on the tree used by Wilf et al.

Black font: shared by Fagus (outgroup) and "Castanoideae". Green font: potential uniquely derived traits. Blue font: traits reconstructed as having evolved in parallel/convergently. Red branches, clades in the used backbone tree that are at odds with currently available molecular data (the N. American relict Notholithocarpus should be sister to the Eurasian Castanea-Castanopsis).

This hardly presents a strong case of crown-group assignation. Except for partial dehiscence, even the modern-day Castanopsis have little discriminating derived traits — they are living fossils with a primitive ('plesiomorphic') character suite. Intriguingly, they are also genetically less derived than other Castanoideae and the oaks (see eg. the ITS tree in Denk & Grimm 2010).

The actual differentiation pattern

The best way to depict what the character set provides as information for placing the fossil is, of course, the Neighbor-net, as shown next.

Neighbor-net based on Wilf et al.'s seven scored morphological traits used to place the fossil. Green: the current molecular-based phylogenetic synopsis — based mostly on Oh & Manos 2008; Manos et al. 2008; Denk & Grimm 2010. I had the opportunity to get familiar with all of the then-available genetic data when harvesting all Fagaceae data from gene banks in 2012 for a talk in Bordeaux. One complication in getting an all-Fagaceae-tree is that plastids, geographically constrained, and nuclear regions tell partly different stories.


Castanopsis, including the fossil, is morphologically a paraphyletic (see also our other posts dealing with paraphyla represented as clades in trees). Note also the long edge-bundle separating the temperate Chrysolepis and chestnuts (Castanea), from their respective cold-intolerant sister genera (Lithocarpus viz Castanopsis) — derived traits have been accumulated in parallel within the "Castanoideae". The scored aspects of Fagaceae morphology are very flexible and ~50 million years is a long time, possibly leading to partial valve indehiscence (or losing it) without being part of the same generic lineage. The puzzling differentiation, and the profoundly primitive appearance of the fossil (shared with modern-day Castanopsis), may in fact be the reason the authors didn't: (i) optimize / discuss very similar, co-eval fossils from the Northern Hemisphere interpreted (and cited) as extinct genera (eg. Crepet & Nixon 1989), (ii) left out the two Fagaceae genera today occurring in South America, (iii) opted for classic parsimony and a partly outdated molecular hypothesis, and (iv) just showed a naked cladogram without branch support values as the result of their "phylogenetic analysis" (Please stop using cladograms!)

Based on the scored characters, the position of the fossil in the graph, and on the background of a more up-to-date molecular-based phylogenetic synopsis (the green tree in the figure above), the most parsimonious interpretation (and probably, the most likely) is that the fossil may indeed be a stem-Castanoideae, a representative of the lineage from which the Laurasian oaks evolved at least 55 million yrs ago (oldest Quercus fossil was found in SE Asia), or even represent a morphologically primitive, extinct (South) American lineage of the Fagaceae. Regarding the "southern route", Ockham's Razor would favor that they are just a South American extension of the widespread Eocene Laurasian Fagaceae / Castanoideae, since very similar fossils and castaneoid pollen is found in equally old and older sites in North America, Greenland (papers cited by Wilf et al.) and Eurasia but not Australia, New Zealand or Antarctica.

A final note: when you have so few characters to compare, you should use OTUs that are not completely ambiguous in every potentially discriminating character, as scored for the "C. fissa group" — the "Castanopsis group" has a single unambiguously defined, potentially derived trait. Using artificial bulk taxa is generally a bad idea when mapping trait evolution onto a molecular backbone tree. Instead, you should compile a representative placeholder taxa set, with as many taxa as you need (or are feasible) to represent all character combinations seen in the modern species/genera.


Postscriptum (14/1/2020)
Relevant matrices (NEXUS-formatted) and explicit character trait maps (Why we want to map trait evolution on networks, pt.1 – Introduction, pt.2 – Topological Ambiguity) have been uploaded to figshare.


Other cited references, with comments
Crepet WL, Nixon KC (1989) Earliest megafossil evidence of Fagaceae: phylogenetic and biogeographic implications. American Journal of Botany 76: 842–855. – introducing a Castanopsis-like infructescence interpreted to represent an extinct genus but very similar to the new Patagonian fossil in its preserved features; and co-occuring with castaneoid pollen (not reported so far for Patagonia) and foliage.
 
Denk T, Grimm GW (2010) The oaks of western Eurasia: traditional classifications and evidence from two nuclear markers. Taxon 59: 351–366. — includes an all-"Quercaceae" ITS-tree (fig. 3) and -network (fig. 4) using data of ~ 1000 ITS accessions; the nuclear-encoded ITS is so far the only comprehensively sampled gene region that gets the genera and main intra-generic lineages apart (recently confirmed and refined by NGS phylogenomic data), something wide-sampled plastid barcodes struggle with. Analysed with up-to-date methods and avoiding long-branch interference by excluding the only partially alignable Fagus, Castanopsis dissolves into a grade in the all-accessions tree and Quercus is deeply nested within the Castanoideae (as already seen in the 2001 tree used by Wilf et al. as backbone). The species-level PBC neighbor-net prefers a ciruclar arrangement in which Notholithocarpus remains a putative sister of substantially divergent and diversified Quercus, followed by Castanea-Castanopsis, and Lithocarpus, while Chrysolepis is recognized as unique.

Oh S-H, Manos PS (2008) Molecular phylogenetics and cupule evolution in Fagaceae as inferred from nuclear CRABS CLAW sequences. Taxon 57: 434–451. – Probably still the best Fagaceae tree, and surely not a bad basis for probabilistic mapping of morphological traits in the family.

Manos PS, Cannon CH, Oh S-H (2008) Phylogenetic relationships and taxonomic status of the paleoendemic Fagaceae of Western North America: recognition of a new genus, Notholithocarpus. Madroño 55: 181–190. – the tree failed to resolve the monophyly of the largest genus, the oaks, but depicts well the data reality when combining ITS with plastid data and, hence, provides a good trade-off guide tree.