Showing posts with label Chinese. Show all posts
Showing posts with label Chinese. Show all posts

Monday, November 26, 2018

How languages lose body parts: once more about structural data in historical linguistics


This is a joint post by Guido Grimm and Johann-Mattis List.

Mattis’ last two blog posts dealt with problems of what linguists call "structural data". Here we discuss what this means for the inference of relationships between languages.

A closer look at structural data: the questionnaire issue

As pointed out before, what is called structural data in comparative linguistics is a very diverse mix of data solely unified by the idea of having some kind of questionnaire that a linguist may use when going into the field and trying to describe a certain language. These questionnaires are a bit different from the traditional concept lists usually used for the purpose of historical language comparison (see the collection of different lists in the Concepticon project by List et al. 2016). The main difference is that they are based on an imaginative question that a field worker asks an informant (which could as well be a written grammar of the language under question). Since questions can be asked in many different ways, while concepts in historical language comparison are usually restricted to the so-called "basic vocabulary", the diversity of structural datasets is much greater than the diversity we encounter when comparing questionnaires based on concept lists.

When analyzing these data, we deal with characters of very different nature, and likely different evolutionary pathways or histories. A biological analogy would probably be (true) total evidence data sets that combine genetic data from: genes/genomes with different inheritance pathways (paternally, maternally, biparentally; basic information level), morphological-anatomical data (visible form, phenotypic), palaeontological data (historical evidence), ontogenetic (life-history stages, developmental features), and biochemical data (expression level). The only difference is probably that the linguistic characters’ histories may be more complex. [Side-remark: ‘total evidence’ datasets found in the biological literature are typically just combination of genetic and morphological data, allowing for the inclusion of extinct/fossil taxa.]

To give a specific example, let's have a look at a the Chinese dataset by Szeto et al. (2018), mentioned in Mattis' blogpost from September. This dataset is now accessible as a GitHub repository (https://github.com/cldf-datasets/szetosinitic). Mattis added some information regarding the different features of the questionnaire. We list these features in slightly abbreviated form in the table below, adding rough categorizations by Mattis in the Comment column.

ID
Description
Comment
p-1
5 or more tone categories
phonological / diachronic
p-2
Retroflex fricative initials
phonological / diachronic
p-3
Bilabial nasal coda
phonological / diachronic
p-4
Stop codas
phonological / diachronic
p-5
Monosyllabic word for 'snake'
lexical
p-6
Differentiation between 'hand' and 'arm'
lexical / semantic
p-7
Differentiation between 'defecate' and 'urinate'
lexical / semantic
p-8
Differentiation between 'eat' and 'drink'
lexical / semantic
p-9
Semantically void suffix in 'table'
lexical
p-10
Different classifiers for humans and pigs
lexical / semantic
p-11
[CLF N] constructions in subject position with definite reference
syntactic
p-12
Reduplicated monosyllabic nouns
morphological
p-13
Post-verbal modal auxiliary developed from 'ge/acquire'
syntactic / diachronic
p-14
Modified-modifier order in animal gender marking
morphological / syntactic
p-15
Post-verbal adverb meaning 'first'
lexical / syntactic
p-16
[V DO IO] order in double object dative constructions
syntactic
p-17
'Give' as a disposal marker
syntactic / diachronic
p-18
'Give' as a passive marker
syntactic / diachronic
p-19
'Go' as a post-VP associated motion marker
syntactic / diachronic
p-20
Marker-Standard-Adjective order in comparatives
syntactic
p-21
case system
morphological / syntactic

Mattis has tried to characterize the features, i.e. matrix’ characters, by generalizing linguistic categories: "phonological", pointing roughly to questions about pronunciation (the biological equivalent would be phenotypic traits in morphology or anatomy); "lexical", pointing to the words in the lexicon (this would be the DNA of a language); "morphological", pointing to the ways in which words are constructed; and "syntactic", pointing to the ways in which words are combined to form sentences. In combination, “morphological” and “syntactic” are equal to ‘meta-level’ biological traits, such as development-related features, ontogenetic evidence, and biochemical composition — the ways in which the genetic code is expressed or used in a living organism in adaption to the environment.

Mattis also flagged some characters as "diachronic", to mark whether the respective feature was selected by the authors due to their independent knowledge about the history of the Chinese dialects. This is something rarely possible in biology, but imagine that we could go back in time to literally observe the evolution of a lineage over a given time-period, and code this observed evolution as traits. Note that this is not entirely science-fiction — there are two examples where we can observe directly pathways of biological evolution: mutation patterns in viruses, and horizontal modification of marine morphs in high-resolution sediment cores.

While one can discuss to what degree a certain feature should belong to this category, it is rather obvious that all phonological features are diachronic, because they name distinctions that reflect well-known processes of sound change, which happened in a couple of Chinese dialects and have been proposed in the past by dialectologists in order to classify the Chinese dialects historically.

For example, consider feature p-3 of the questionnaire: Does a given dialect have a syllable that ends in [-m]? From the history of the Chinese dialects we know that the [-m] was present in Middle Chinese, but later merged with [-n] and [] in many varieties. Given that we know that this happened, and that we know that people have used this to mark a split, especially between the "innovative" dialects in the North and the South, it is clear that this feature bears explicit historical information. The same holds for all phonological features that we find in the data: p-1, the number of different tones in the dialects is again roughly reflecting the differences between languages in the North and in the South (the North having lost many tones); p-2 reflects the retention or specific development of retroflex sounds (similar to sh in English as opposed to s) mostly in the North; and p-4 reflects if a variety has syllables that can end in [-p, -t, -k], again a feature characteristic for the more "conservative" varieties in the South of China.

Figure 1: Overlap of features in Szeto et al.'s (2018) structural feature collection of Chinese dialects

Four lexical features have further been flagged as "semantic"; we query here existing or missing distinctions of concepts. People who learned, for example, Russian or certain German dialects know that it is rather common to have a single word for what other languages call "arm" and "hand" (see the respective entry in the CLICS database) or "foot" and "leg".

This diverse feature collection is coded as binary characters, reflected by presence/absence, or a yes/no answer to the question in the questionnaire. The choice of features is very selective. A biological analogy would be a matrix collecting incompatible splits of paternal (molecular) genealogies, along with a few prominent phenotypical traits (reflecting major evolutionary steps), and some traits that we expect to be primarily triggered not by genetics (inheritance) but by expression or adaptation to the environment. Biologists would not phylogenetically analyze such diverse and complex, potentially selection-biased data (although it could be very interesting), but linguists do.

In this context, it is remarkable, but also typical for these kind of data, that the 21-character feature collection by Szeto et al. (2018) has no feature in common with the collection by Norman (2003), a 15-character-matrix, which we also converted to our Cross-Linguistic Data Formats (see Forkel et al. 2018) in order to increase the data comparability.


Figure 2: A Neighbor-net splits graph of the structural data by Szeto et al. (2018).
The typification, coded as binary matrix to infer the Neighbor-net splits graph in Figure 2, demonstrates some basic characteristics of such 2-dimensional graphs. Note four of the 'characters' (typification categories) correlate with an edge(-bundle) in the network, separating the 'taxa' (the queried features). All "semantic" taxa are also "lexical", but "lexical" is more comprehensive, hence, "semantic" is placed as 'descendant' of "lexical" (Neighbor-nets can visualize ancestor-descendant relationships to some degree). "Morphological" taxa are either just "morphological" or also "syntactic", hence the pronounced box.

For "diachronic" and "syntactic", we have no corresponding edge(-bundle), because one taxon is also "lexical", but the others are "diachronic" and "syntactic" — this is a conflict that cannot be resolved with two dimensions. To visualize all the resultant 'taxon' splits, called also taxon bipartitions, we would need a third dimension. Lacking a third dimension, the Neighbor-net prioritizes keeping most "syntactic" together, because the "diachronic-syntactic" are closer to "syntactic" (max. 1 'character' difference) than to "diachronic-phonological" (2 character difference). The "syntactic-lexical" has to be placed apart because it is equally close to "lexical" and "syntactic" 'taxa', but differs much from "morphological-syntactic" or "diachronic-syntactic", the closest two relatives of "syntactic"-only 'taxa'. It is resolved closer to the centre of the graph, because it is more closely related to the other "syntactic" taxa than to the rest of the "lexical" taxa. This is also the reason why the "syntactic"-only taxa have to be placed farther out: "Diachronic-phonological" and "syntactic-lexical" are closer to the other endpoints, and the distance of "syntactic"-only to "diachronic-phonological", "lexical" and "morphological" should be as large as possible.

Losing body parts: How data coding masks underlying processes

Most typologists collecting structural data are not per se interested in phylogenies. Yet, given that scholars deliberately collect historical (diachronic) features, this shows that even if they would not necessarily admit it, they have a genuine interest in uncovering the history of the languages under question; or at least, how closely related languages (or here: dialects) are. But this requires understanding the characters we analyze, the collected "structural data".

In evolutionary biology, the key question people (should) ask when trying to select characters is how their change can be modeled on a tree or a network. What processes could be expected that shaped the data? What is behind the diversity? Is similarity or dissimilarity instigated by:
  • [A] inheritance, i.e. passed from an ancestor to all / some of its descendants,
  • [B] random mutation and/or sorting, i.e. the product of a stochastic, evolutionary neutral process,
  • [C] non-random mutation, i.e. processes that recur frequently, may be beneficial and positively (gain, or negatively: loss) selected for, or
  • [D] secondary contact, mixing of lineages by hybridization (symmetric mixing) and introgression (asymmetric mixing)?
[A]–[C] are vertical processes following a tree, even if the tree does not necessarily need to be the same; [D] is (mostly) horizontal and can only be modeled using a network. For each of the above, we can find an analogy in the evolution of languages.

In addition, process [3], and to a lesser extent [4], can lead to what biologists call 'homoplasy', meaning that the same feature is observed in two unrelated or distantly related taxa. In the context of phylogenetic inferences, homoplasies inflict tree-incompatible signals, seemingly reticulate patterns originating from a tree-like evolution. Structural (or other) linguistic data and phenotypical biological data have a lot in common — complex processes are boiled down to mere absence or presence of features (or traits, as they are called in biology).

Figure 3: Basic evolutionary processes, we need to consider when looking at linguistic data. Or biological traits, when we replace simplification by adaptive evolution, positively selected traits.

If we check the features in our table above, and ask: to which degree can they be used to model these processes (see also David's last post on illogic in phylogenetics), e.g. simply distinguish between similarity by chance, relatedness, or secondary contact (mixing), we can easily see that they are by no means optimal for evolutionary investigations. This is not necessarily because of the processes they involve, but more importantly because of the data sampling, which makes modeling almost impossible, with each character needing its own model.

As an example, take the feature p-6 in our table. Whether or not a language makes a distinction between "arm" and "hand" does not seem to follow specific geographic or genealogical patterns. The following figure shows a plot from the CLICS database (List et al. 2018), visualizing the most frequently recurring polysemies (or colexifications) centering around the concept "arm". The full visualization in CLICS can be found here, and when hovering with the mouse over the link between "arm" and "hand" (marked in green below).

Figure 4: Colexification network in the CLICS database.

From eye-balling the data, it is hard to find a consistent geographic / language-family pattern, which suggests that the feature p-6 is likely to show a high degree of homoplasy in the languages of the world. Obviously, different people decided not to distinguish between "hand" or "arm". But, the example of the Sami languages in northern Scandinavia also demonstrate that some people using related, long-isolated languages, consistently don't make the distinction. Here, the homoplasy is inherited (lineage-conserved). A biological analogy would be the rarely applied difference between a 'convergence' (a trait is independently evolved in different lineages) and a 'parallelism' (a trait is expressed by different but not all members of the same lineage).

Figure 5: Geographic distribution of arm/hand colexifications in the CLICS database.

A specific analogy to the "hand-arm" colexification / differentiation pattern is leaf shedding in oaks and their relatives (Fagaceae, the beech family). Some oak lineages (section Cerris of oaks, beech trees, chestnuts) are essentially or strictly deciduous, others (sections Cylcobalanopsis, Ilex, the sister sections of Cerris; Castanopsis, the sister genus of chestnuts) are always evergreen, and the biggest group (number of species) of all Fagaceae, subgenus Quercus includes evergreen (1 section), mixed (the two by far largest sections), and deciduous (1 nearly extinct section) sublineages. To some extent this is linked to the climate in which the species thrive (high latitudes and/or per-humid = deciduous, low latitude and/or seasonally dry = evergreen), but consistently evergreen and deciduous lineages do co-exist.

Looking at the Chinese dialects, we see that p-6 represents a trivial split in the network.

Figure 6: A Neighbor-net inferred from the Szeto et al. matrix. Dialects that distinguish "arm" and "hand" with filled dots ('1' for character 6 in the matrix), those that don't ('0') with empty dots. We can put a single line separating all don't- from do-taxa (dialects), i.e. a bipartition of the taxon set fitting the character partition seen in (p-)6.

But, given the general patterning of the feature on a global scale, does this really mean that it is inherited — that is, a good feature to reflect relatedness?

Whether a feature is likely to be homoplastic is just one part of the story. Linguists typically have more information about how things change than do biologists, putting a double-edged sword in their hands (that they hardly ever use). Asking whether "hand" and "arm" are expressed by distinctive concepts does not consider the underlying processes. Here, we can assume at least three different character states, namely:
  1. "arm" and "hand" are expressed by the same word, which is the original word for "arm",
  2. "arm" and "hand" are expressed by the same word, which is the original word for "hand", and
  3. "arm" and "hand" are expressed by different word.
We could even have a forth state, in which "arm" and "hand", in the whole long history of the ancestral languages, was always used to express "arm or hand" (i.e., both body parts). No differentiation and no later generalization from either arm nor hand took place.

Figure 7: Left, current scoring; right, scoring taking into account the actual mutation process.

From Ancient Chinese, we know that "1" (Yes, I do differ between "arm" and "hand") was most likely the original state. We can further assume that once the distinction is dropped, it is less likely to come back again (although this can, of course, also happen). That is, our model involves two possible mutations (vertical process): we lose the word for "arm" due to its replacement by "hand", or we lose the word for "hand" due to its replacement by "arm", each with its own probability.

Figure 8: Probability distribution for transitions involving "hand" and "arm".

The probability, mutation or not, and which mutation, relates to four principal driving factors:
  1. probability of random loss (mutation)
  2. probability of random gain (mutation)
  3. global linguistic tendencies
  4. regional socially-enforced preference
Establishing p-arm (loss "arm") and p-hand (loss "hand") is not trivial, because they may be affected by what is the word for "arm" and "hand" (for simplicity we will assume that p+arm and p+hand are close to 0). We could expect a higher tendency to keep the word that is easier to pronounce or less easy to confuse with other words and, hence, is easier to understand. If two dialects with different states come into contact, this may also influence the decision to take over a state or not. In everyday language, a distinction between "arm" and "and" may be useless because of the clear context in which both words are used, so p1-word > p2-words. However, closeness to administration centers or areas with a higher percentage of educated people could decrease p1-word, because it may be considered a sign of poor social standard to not make the difference between "arm" and "hand".

Figure 9: Vertical and horizontal processes involving transitions of "hand" and "arm".

Estimating p can only be left to phylogenetic algorithms (unless more detailed information is available). But we can (and should) design the questionnaire to capture as many of the processes as possible. In this case, to not only ask whether there is a distinction between "arm" and "hand", but also to find out whether the word "arm" or "hand" is used, e.g. by using two questions/binary characters:
  • Do we use "hand"?
  • Do we use "arm"?
Note that this question requires quite a deal of knowledge about the languages under investigation, since it may not be trivial to find out what was the "original" word for "arm" or "hand".

Therefore, a further step would be to replace the binary characters by a value measuring the similarity between the words used for "hand" and those used for "arm". One could again argue that adding this information would add historical information to the feature, but it is clear that the abstract nature of the question is hiding important phylogenetic (and also typological) information from us.

It seems therefore, that, instead of asking whether or not there is a distinction between "arm" and "hand", it would make much more sense to trace the cognacy (or homology) of the expressions for "arm" and "hand" across all taxa (languages, dialects), and think of ways how this could be scored and modeled by phylogenetic analyses. The structural data framework with its features based on simple yes-no questions therefore inevitably leads to a misinterpetation of processes when analyzing the data with phylogenetic software.

The need for exploratory data analysis

In reality, structural (or other) data sets in linguistics face problems similar to the ones palaeontologists face when trying to establish phylogenetic relationships between fossils (extinct organisms) — the probability for a mutation (visible change) is largely unknown, and differs not only from character to character but also within the same characters. A state 0, 1, 2 etc. may have a higher probability to manifest (or get lost) in one lineage than in another.

In addition, the linguistic problems recur in a similar way to that of biologists working close to and below the species level (see also Guido's post on population dynamics and individual-based fossil phylogenies) — reticulation is rather the rule than the exception, as similarity is triggered by contact,  so that horizontal processes, not inheritance, may dominate evolutionary dynamics. Thus, the diversity pattern cannot be modeled by a tree alone. Establishing explicit probabilistic frameworks to deal with this may not only be difficult but even impossible (given the available data). Meanwhile, however, one can embrace exploratory data analysis as a heuristic tool.

So, let's look at the example. As in the original paper, we used the binary matrix of the 21 characters to infer a planar, 2-dimensional (meta-)phylogenetic network, a Neighbor-net splits graph. The resulting graph is a longitudinally inflated spider-web, with its endpoints defined by the southern Chinese dialects (e.g. Guangzhou, Nanning, Taishan) and the north-central (eg. Linxia and Xining) dialects. The latter are significantly closer (geographically and data-wise) to the Bejing version of Chinese.

Figure 10: The Neighbor-net based on simple mean (Hamming) pairwise binary character distances

The first thing to note is that the matrix includes dialects that are indistinct (green stars) for all 21 characters, and some that are geographically and data-wise very similar to each other, while being distinct from all others (green ovals). In biology, we call this (taxic, lineage-)coherence. In addition to Linxia and Xining, we have Nanchang and Lichuan characterized by elongated ('tree-like') terminal edge-bundles. These obviously represent closely related dialects sharing a long(er) common history.

Others have more than one possible closest relative. For instance, Liuzhou may share quite a few features with Guangzhou, but it is equally close to the Nanchang-Lichuan pair (yellow fields). Dongtai (orange star) is unique, but its 'neighborhood' (orange-ish brackets) as defined by shared edge-bundles that include Changsha (which again is most related to Jiujang) and Taiyuan plus Baotou, the latter two substantially closer to the Bejing (red star) group.

Similar to Dongtai, and also connected to the central part of the graph, are dialects with long-terminal branches (edges). Hefeng (blue star) is substantially different from Dongtai, and only has one further dialect in its neighborhood (blue bracket), Wangrong, a close relative of the Bejing group. The Wuhan, Chengdu, and Guiyang (gray field) dialects appear, on the other hand, to be completely isolated.

As explained above, there are different processes, vertical and horizontal ones, that may trigger similarity, and we want to get an idea as to which character may be influenced by which process. From the graph, several aspects are obvious:
  • geographic closeness plays a major role,
  • the signal provided by the data is not tree-like,
  • the data is highly homoplastic, and includes internal conflict.
Not so obvious is whether this situation is due to random or evolutionary directed similarity, or reticulation. Since the graph is planar, and puts the Chinese dialects in a circular order, we can order the character matrix accordingly to see how the traits form groups (which could be called cliques in this context). In the next step, we can then map each character onto this network, to see how well they fit with the overall similarity pattern. We showed this above for p-6 (hand-arm-distinction, one split), and here we add a character with quite a poor fit, p-17 (syntactic-diachronic), "give" as a disposal marker.

Figure 11: Character mapping for p-17 (filled dots, "give" used as disposal marker; empty, not used), with the p-6 split indicated as well. Red, splits (taxon bipartitions defined by character cliques) that have no corresponding edge-bundle (neighborhood); blue, splits with neighborhood; green, unique, isolated change (deviation from the rule) within the neighborhood.

The number of inferred mutations in the map uses Ockham’s Razor, upon which parsimony (tree and network) inference relies as well. Using such a map, we can even provide an estimate for how likely (qualitatively spoken) a change is under the assumption that neighborhoods in the graph represent either exchange (homogenization) between closely related dialects or are inherited, reflecting both horizontal and vertical relatedness. Mapping characters on a 2-dimensional network allows finding a scenario beyond a single tree hypothesis.

For p-6, we need just one change (i.e. loss in all more south-bound dialects), but we don't find an edge bundle corresponding to this unique change. Given what we discussed above about p-6, we have more independent losses than the simple reconstructed one. Social preference or general contact for retaining the primitive state of having two words could explain why dialects closer to the Beijing dialect area have a "0", although not all are closely related in general.

For p-17, we need at least four (independent) changes from "0" → "1", two of which have a corresponding edge bundle (blue, Nanchang plus Lichuan, Changsha plus Dongtai), one isolated (green, Luoyang), and one without a corresponding edge bundle (Wuhan and Hefeng dialects). The (equally parsimonious) alternative for p-17 would be a series of gains and losses, with the same number of steps:

Figure 12: Alternative scenario for p-17.

This is where one needs to consider additional knowledge about the probability of getting or retaining a certain feature. The state shared by most dialects across the entire net is “0”, irrespective of overall similarity, which would make it a natural pick for the primitive state. Thus, assuming four (or more) changes from 0 → 1 (acquisition of the queried feature), rather than two independent acquisitions (starting with the Beijing group; note, the position of the root will not change the number of needed changes), then a loss (1 → 0) in many southbound dialects and a re-gain (0 → 1) in the Nanchang + Lichuan dialects.

The same assessment can be made for all of the characters, and we end up with something like this:

Figure 13: Fully annotated split network of the data. Changes relating to edge-bundles accordingly colored, arc indicate changes without a corresponding edge-bundle. Note, the prominent yellow split that defines a neighborhood of dialects most similar to the Beijing dialect, albeit there is no character supporting this edge. The rather poor fit of many character splits (cliques) with edge-bundles relate to the fact that we visualize a highly complex diversification (multi-dimensional processes) using a planar, 2-dimensional graph.

While this figure may be confusing at first sight, it comprehensively shows what the characters contribute to the overall graph. We can discriminate more-likely from less-likely mutations (how many changes are needed at least), but also the character assemblies shared by putatively closely related dialects.
  • p-3 and p-11 are a typical feature of Guangzhou and allied dialects within the southern Chinese complex. p-3 is also present in Lichuan, and p-11 in Jixi (thus in not so distant dialects).
  • Features p-6 to p-9, p-16, and p-19 form a diagnostic suite for the Guangzhou dialects and other dialects related to them in the one or other fashion and distinguish them from, e.g., the Beijing group
  • The latter, the Beijing group, has fewer diagnostic character assemblies. One characteristic sequence could be p-1, p-2, p-12, p-14, but this includes three features with a minimum of 3+ changes. Similarity here is mostly the result of a lack of (potentially) derived features (hence, the character-unsupported yellow edge-bundle defining a Beijng-including neighborhood)

Outlook and summary

In this re-investigation, we have, once more, commented on the problems we see with the use of structural features for the purpose of historical language comparison and phylogonetic reconstruction. We see the major problems in the (often) unfortunate choice of question, resulting in elicitations of features that cannot be easily modeled with current software for phylogenetic analyses. It is important to keep in mind, in linguistics and phylogenetics, that we can infer trees or networks based on data of no matter what quality and information content. But before we present the result, we should have taken a look at the primary data.
  • Does it fit with the resulting graph, or not?
  • Where does it fit, and where not?
In the context of our critique of linguistic questionnaires, the mapping strategy discussed above opens a potential avenue to identify:
  • stable / unstable features (geographically or evolution-wise) and
  • coherent / incoherent features.
Based on this, we can then inquire as to which degree language (or dialect) groups influenced, stabilized or modified each other by geographic proximity.

Inference-wise, the natural next step would be to use the information about the minimum number of necessary changes to counter-weight characters. This would eventually allow to use median networks (and related) approaches on the data, which is currently the only way to explicitly identify ancestors using phylogenetic reconstructions. With the current matrices, the extreme homoplasy makes an unweighted application of median networks and related methods impossible.

References

Forkel, R., J.-M. List, S. Greenhill, C. Rzymski, S. Bank, M. Cysouw, H. Hammarström, M. Haspelmath, G. Kaiping, and R. Gray (2018) Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics. Scientific Data 5.180205: 1-10.

List, J.-M., M. Cysouw, and R. Forkel (2016) Concepticon. A resource for the linking of concept lists. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation, pp. 2393-2400.

List, J.-M., M. Walworth, S. Greenhill, T. Tresoldi, and R. Forkel (2018) Sequence comparison in computational historical linguistics. Journal of Language Evolution 3.2: 130–144.

Norman, J. (2003) The Chinese dialects. Phonology. In: Thurgood, G. and R. LaPolla (eds.): The Sino-Tibetan languages. Routledge: London and New York, pp. 72-83.

Szeto, P., U. Ansaldo, and S. Matthews (2018) Typological variation across Mandarin dialects: An areal perspective with a quantitative approach. Linguistic Typology 22.2: 233-275.

Supplementary data

The data we used to create the analyses and figures provided in this post are available at https://github.com/cldf-datasets/szetosinitic/tree/master/examples

Monday, May 28, 2018

Comparing reconstruction systems in historical linguistics


The term linguistic reconstruction has a very specific meaning in historical linguistics, pointing usually to the techniques that are used in order to infer how a given language was originally pronounced, even though it has not been attested in written sources. In previous posts, I have occasionally pointed to reconstructed forms, the so-called proto-forms, which linguists usually mark as such by putting an asterisk in from of them. For example, the word Indo-European *ph₂tér- is a reconstructed proto-form for the supposed Indo-European word "father".

While the reconstruction techniques are usually limited to languages for which we have no written record, they can in principle also be applied in order to find out how ancient languages like, for example, Latin and Greek, were pronounced in detail (Sturtevant 1920). For languages like Chinese, whose writing system leaves almost no clues about pronunciation, linguistic reconstruction is the only way to investigate the pronunciation of the oldest stages of the language.

When dealing with different reconstruction systems for Old Chinese phonology, it is quite difficult, even for experienced scholars, to spot the actual differences between the systems. That these differences exist, and that they can be quite substantial, is beyond question — and easy to understand, if one takes into account that Old Chinese is reconstructed with the help of a philological (as opposed to a mainly comparative) approach, by which data from different sources is sifted and individually weighed (cf. Jarceva 1990: 409 and List 2008).

When comparing different reconstruction systems, it is not enough simply to look at the inventories of proto-phonemes proposed by different scholars. Even if two proto-inventories (the sets of the reconstructed sounds) are exactly the same, it is possible that scholars will provide different reconstructions for individual characters. The only way to compare two or more reconstruction systems is therefore to compare the concrete reconstructions for a certain number of characters.

In addition to the sample of words, however, we also need a clear account of which segments (which proto-sounds) should be compared with each other. When comparing proto-forms for Chinese 一 ‘one’ in different Old Chinese reconstruction systems, such as Karlgren (1950) *ʔi ̯ĕt, Li (1971) *ʔjit, Wáng (1980) *iet, and Baxter and Sagart (2014) *ʔi[t], we would obviously not compare the medial *i ̯ of Karlgren with the initial *ʔ of Baxter and Sagart.

When adding more reconstructions, such as the one for 七 ‘seven’ across the four systems, for which the authors give *ts'i ̯ĕt, *tshjit, *tshiet, and *[tsh]i[t], respectively, we can further see that there are not only differences for the different segments in the same positions, but also for the interpretation of the words. Although all authors give different medials, main vowels, and finals in the two words, they are structurally consistent in giving both words the same sound segments for medial, nucleus, and coda.

What we can see from this example is that any difference in the sound segments, like the choice of initials, or the concrete solution proposed for a problem, does not immediately reflect important differences in the reconstruction systems. If two scholars just choose another symbol for a distinction that they both recognize and acknowledge, this does not render the reconstructions incompatible. It should therefore not be used as a criterion for dismissing a given reconstruction system, at least not in a first step. If two systems are structurally equivalent, then they have equivalent predictive power for the descendant language(s) they are supposed to reconstruct.

This abstractionist notion of proto-forms, which can be found in the early work of Saussure (1916) and Meillet (1903), is problematic for the endeavour of linguistic reconstruction, and usually not strictly followed (Lass 2017). Nevertheless, the potentially abstract notion of proto-forms is important to be kept in mind when comparing different reconstruction systems. When distinguishing the structural differences (which result from the direct interpretation of the data and the identification of regular sound correspondences) from the substantial differences (resulting from a phonetic and phonological interpretation of the identified correspondences), we have a much clearer account of the core of the differences, and whether they are worth our consideration or not.

But how can we compare reconstruction systems structurally? Firstly, we need to have the data assembled in aligned form, in order to make sure that we only compare like with like (e.g., medial with medial, and vowel with vowel). A sample illustration in which alignments of the proto-forms for ‘seven’ and ‘one’, produced with the help of the EDICTOR tool (List 2017), is given in the figure below.

Comparing reconstruction proposals with the help of alignments.

Alternatively, we can also select a single aspect, such as, for example, the vowel system proposed in different reconstruction systems. Having assembled a substantial amount of different proto-forms in this way, the structural comparison between different reconstruction systems can be modeled as a comparison of different cluster analyses, or, more accurately, partitioning analyses. A partitioning analysis assigns a given number of objects to a certain number of different groups. When dealing only with the vowels proposed by different reconstruction systems, we can say that a given reconstruction, like the one by Karlgren, for example, assigns each Chinese character, for which a proto-form is given, to a particular group depending on the main vowel selected for the reconstruction.

If, for a given number of reconstructions, we model each reconstruction system as a partitioning analysis, based on the main vowel proposed by the system, we can use standard metrics from graph theory and Natural Language Processing to compare different reconstruction systems with each other. Very straight-forward measures for the comparison of two partitioning analyses are the so-called B-Cubed scores (Amigó et al. 2009), which have proven specifically useful for the evaluation of automatic cognate detection methods in historical linguistics, compared to a gold standard (Hauer and Kondrak 2011, List et al. 2017).

Being an evaluation measure, B-Cubed scores come in the typical three flavors of precision, recall, and F-Score. Precision is similar to the notion of true positives, and recall is similar to true negatives. For the purpose of comparing reconstruction systems, only the F-score is needed, as it is a symmetric measure, and the notion of true positives and true negatives is meaningless, unless we decide that we blindly trust one of the given systems. As also for the scores for precision and recall, the F-score ranges between 0 and 1, with 1 indicating that the two partitioning analyses are identical.

In order to compare more than one reconstruction system, we can make use of techniques for exploratory data analysis (Morrison 2014); and the most straightforward way to do this, is, of course, to use the NeighborNet algorithm (Bryant and Moulton 2004), as provided by the SplitsTree package (Huson 1998).

In order to illustrate how data-display networks can be used to study differences among Old Chinese reconstruction systems, I designed a little experiment, based on data taken from (List et al. 2017b), who provide Old Chinese reconstructions for all rhyme words in the Shījīng based on eight different reconstruction systems (Baxter and Sagart 2014, Karlgren 1950, Li 1971, Pān 2000, Schuessler 2007, Starostin 1989, Wáng 1980, Zhèngzhāng 2003).

In order to keep the analysis simple, I extracted only the different reconstructions of the main vowel for each character in each system, and carried out a pairwise comparison of all eight systems, computing the B-Cubed F-scores for each pair, omitting characters for which no reconstruction could be found in the data. These scores were then converted to a distance matrix, and fed to the NeighborNet algorithm (the source code can be downloaded here). The resulting network is provided in the figure below.

NeighorNet reflecting the closeness of the different reconstruction systems
As one can see, the data roughly clusters into three subgroups, namely Schuessler, Baxter and Sagart, and Starostin vs. Pān and Zhèngzhāng vs. Karlgren, Li, and Wáng. On a larger scale, we can divide the data into all six-vowel systems versus the non-six-vowel systems (Karlgren, Wáng, Li). Given that Pān is a direct student of Zhèngzhāng, the closeness between their reconstruction systems is not surprising.

What may be surprising is the closeness of the Schuessler, Starostin, and Baxter and Sagart systems, given their notable differences with respect to the criterion of vowel purity tested by List et al. (2017b). Even if the network analysis cannot directly explain all of these differences in detail, it seems like a worthwhile enterprise, which should be further expanded by comparing not only the vowels, but fully aligned proto-forms.

Given the straightforwardness of the application, it seems also useful to test it on other language families where there is similar disagreement, as in the reconstruction of Old Chinese phonology.

References

Amigó, E., J. Gonzalo, J. Artiles, and F. Verdejo (2009): A comparison of extrinsic clustering evaluation metrics based on formal constraints. Information Retrieval 12.4. 461-486.

Baxter, W. and L. Sagart (2014) Old Chinese: a new reconstruction. Oxford University Press: Oxford.

Bryant, D. and V. Moulton (2004) Neighbor-Net. An agglomerative method for the construction of phylogenetic networks. Molecular Biology and Evolution 21.2. 255-265.

Hauer, B. and G. Kondrak (2011) Clustering semantically equivalent words into cognate sets in multilingual lists. In: Proceedings of the 5th International Joint Conference on Natural Language Processing. AFNLP 865-873.

Huson, D. (1998) SplitsTree: analyzing and visualizing evolutionary data. Bioinformatics 14.1. 68-73.

Jarceva, V. (1990) Sovetskaja Enciklopedija: Moscow.

Karlgren, B. (1950) The Book of Odes. Chinese text, transcription and translation. Museum of Far Eastern Antiquities: Stockholm.

Lass, R. (2017) Reality in a soft science: the metaphonology of historical reconstruction. Papers in Historical Phonology 2.1. 152-163.

Li Fang-kuei 李方桂 (1971) Shànggǔyīn yánjiū 上古音研究 [Studies on Archaic Chinese phonology]. Qīnghuá Xuébào 清華學報 9.1-2. 1-60.

List, J.-M. (2008) Rekonstruktion der Aussprache des Mittel- und Altchinesischen. Vergleich der Rekonstruktionsmethoden der indogermanischen und der chinesischen Sprachwissenschaft [Reconstruction of the pronunciation of Middle and Old Chinese. Comparison of reconstruction methods in Indo-European and Chinese linguistics]. Magister thesis. Freie Universität Berlin: Berlin.

List, J.-M., S. Greenhill, and R. Gray (2017) The potential of automatic word comparison for historical linguistics. PLOS One 12.1. 1-18.

List, J.-M. (2017) A web-based interactive tool for creating, inspecting, editing, and publishing etymological datasets. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics. System Demonstrations. 9-12.

List, J.-M., J. Pathmanathan, N. Hill, E. Bapteste, and P. Lopez (2017) Vowel purity and rhyme evidence in Old Chinese reconstruction. Lingua Sinica 3.1. 1-17.

Meillet, A. (1903) Introduction à l’étude comparative des langues indo-européennes. Hachette: Paris.

Morrison, D.A. (2014) Phylogenetic networks: a new form of multivariate data summary for data mining and exploratory data analysis. WIREs Data Mining and Knowledge Discovery 4: 296-312.

Pān Wùyún 潘悟云 (2000) Hànyǔ lìshǐ yīnyùnxué 汉语历史音韵学 [Chinese historical phonology]. Shànghǎi Jiàoyù 上海教育: Shànghǎi 上海.

de Saussure, F. (1916) Cours de linguistique générale. Payot: Lausanne.

Schuessler, A. (2007) ABC Etymological dictionary of Old Chinese. University of Hawai’i Press: Honolulu.

Starostin, S. (1989) Sravnitel’no-istoričeskoe jazykoznanie i leksikostatistika [Comparative-historical linguistics and lexicostatistics]. In: Kullanda, S., J. Longinov, A. Militarev, E. Nosenko, and V. Shnirel’man (eds.): Lingvističeskaja rekonstrukcija i drevnejšaja istorija VostokaMaterialy k diskussijam na konferencii.[Materials for the discussion on the conference].1. Institut Vostokovedenija: Moscow. 3-39.

Sturtevant, E. (1920) The pronunciation of Greek and Latin. University of Chicago Press: Chicago.

Zhèngzhāng Shàngfāng 郑张尚芳 (2003) Shànggǔ yīnxì 上古音系 [Old Chinese phonology]. Shànghǎi Jiàoyù 上海教育: Shànghǎi 上海.

Monday, January 29, 2018

Networks of pronunciation glosses in Traditional Chinese phonology


Every person who has learned how to read in any language will at some point have to deal with the question of how to pronounce words with unusual spellings. The English writing system offers an abundance of examples; and in my own pronunciation practice of English, I am still frequently corrected by native speakers when mispronouncing words that I know only from books.

For example, what constitutes a big problem for me is the stress on words of Latin origin, which have different stress patterns in German, my native tongue. While we speak of a Theo·rem in German, stressing the final syllable, English pronounces the word as the·orem, stressing the first. But my problems with the English writing system (and probably the problems of many other non-native and native speakers as well) do not usually end with the placement of stress, but may often go much deeper.
Pronunciation

Determining how an unknown word is pronounced in English is very easy nowadays. One can just use one of the numerous online dictionaries, where pronunciations are given in form of sound files. Another, more old-fashioned, alternative is to consult a classical dictionary that illustrate pronunciation with help of the International Phonetic Alphabet (IPA 1999). As linguists, we use it on a regular basis in order to compare pronunciations of words across different languages and language families. The original purpose of the IPA, however, was essentially the correct pronunciation for first and second language acquisition; and many teachers were involved in its creation in the late 19th century (compare Kalusky 2017).

Even earlier than the standardization efforts by the International Phonetic Association are ad hoc practices of glossing the pronunciation of difficult words, by comparing them with the pronunciation of more common words in the same language. In order to explain the pronunciation of the English word digest, for example, we could say that word is pronounced as die in dead and as gest in adventure. In the English context, it seems furthermore to be common to make use of some very basic syllables that most people will read and pronounce unambiguously, like ah for the a we find in abacus as opposed the the a we find in and, or toe for the "normal" o-sound we find in no as opposed to the sound of the o in words like to.

That these ad hoc systems, which humans use to gloss pronunciations in writing, are not very reliable can be easily understood when recalling that writing systems have often grown over centuries, reflecting different layers of pronunciation practices applied to words that were imported into the languages at different stages in history. English is, of course, an extremely messy case, but even writing systems like German or Russian, of which speakers would say that the pronunciation is close to the spelling, are far from reaching the explicitness of the International Phonetic Alphabet.


Chinese pronunciation

A particularly interesting case concerns historical glossing practices in the history of Chinese. As I mentioned in an earlier blogpost on networks in Chinese poetry, the Chinese writing system gives only minimal hints regarding the pronunciation of its characters. A character like 手 "hand", which is pronounced as shǒu (or [ʂɔu²¹⁴] in the IPA), does not tell us anything about its pronunciation; and even its meaning is difficult to derive from its modern written form.

Chinese scholars became aware of the problem rather early, around the 1st century AD when they tried to read the ancient texts produced by their intellectual and poetic masters some 500 years before. In order to make sure that the pronunciation of infrequent characters would not be forgotten over time, they developed different ways to gloss character pronunciations in a more or less systematic manner.

The ancient Chinese scholars didn't have an alphabet to simply transcribe their sounds — intensive contact with Indian phoneticians started much later. So, they started from simple equations, according to which one character was pronounced similarly to another character.

For example, the Shuōwén Jiězì (Explaining Simple and Complex Characters) is an early Chinese character dictionary by the famous scholar Xǔ Shèn (58-148 AD), which was published in 121 AD. In it, the author occasionally uses the formula "read [this character] as X" (in Chinese 读若 dúruò X), in addition to his explanations of the meanings and the structure of the characters. The disadvantage of this duruo method, as linguists often call it (Coblin 1983), is that it only allows glossing of characters for which a simple character with an identical pronunciation exists. It is also not clear whether the formula consistently points to strictly identical pronunciations or whether certain deviations are allowed.

In order to overcome these problems, much more precise ways of glossing character pronunciations were developed from about the 2nd century AD. One of the most interesting glossing systems in this context is the so-called fǎnqiè spellings (Coblin 1983, Branner 2000). This spelling method, which seems to go back to at least the third century, is based on breaking the character pronunciation into two parts, the initial and the final, and selecting one character for glossing each of the two parts — one with an identical initial sound and one with the identical final. If we applied this method to English, we could think of explaining the pronunciation of rice as rye-nice, with rye pointing to the initial sound r and nice pointing to the final of the word.

In the following figure, I have tried to illustrate how both methods (the dúruò and the fǎnqiè method) are applied in concrete examples of text.


Given their straightforwardness and simplicity, fǎnqiè pronunciation glosses became quite popular among Chinese scholars. Even today, people may occasionally use them in order to explain pronunciations without having to rely on foreign writing systems, like the Latin alphabet. As a result, there is an abundance of sources that use this pronunciation device throughout the history of the Chinese language. Although the pronunciation is only given indirectly, with respect to the pronunciation traditions that were active during a given epoch, the fǎnqiè spellings offer great help to explore how the pronunciation of the Chinese language changed over time.

Pronunciation networks

Most of this research on the usage of fǎnqiè spellings has been carried out manually. The first work on fǎnqiè spelling goes back to the early 19th century, when scholars like Chén Lǐ (1818-1882) began to investigate systematically which characters were used to denote certain initial sounds (in Chinese, these are called the upper fǎnqiè characters, fǎnqiè shàngzì 反切上字), and which characters were used to denote the finals (called the lower fǎnqiè characters, fǎnqiè xiàzì 反切下字).

As we might expect, instead of using the same character for the pronunciation of the initial sound all the time, scholars would often alternate the characters, but the alternations were more or less consistent, with some characters being used more frequently and some characters being used less frequently. Scholars like Chén Lǐ figured out that the characters could be classified in a rather rigorous manner which would allow us to reconstruct direct pronunciations of the fǎnqiè spellings.

For example, based on the spellings reported in the Qièyùn, an early rhyme book published in 601 AD, we can say that the characters gōng 公, 古, gàn 干, etc. were regularly used to indicate initials that would be spelled as [k] in the International Phonetic Alphabet, while kǒu 口, 可, and 苦 were used to pronounce [] (a k with strong aspiration).

What I find even more interesting and important than these concrete findings, is that Chinese scholars inherently employed rudimentary network thinking to arrive at their clusters (Gēng 2004). The system of glossed character and glossing character can be easily translated into a system of directed networks, in which we draw a link from the glossing character to the glossed character.

For a talk held earlier during the last year (List 2017), I constructed such a network from the Guǎngyùn (ca. 1000 AD), a later edition of the aforementioned rhymebook Qièyùn, which gives fǎnqiè spellings for more than 20,000 characters. In this network, I concentrated only on the initials, that is, the initial consonants of the language encoded in the source, and constructed a network of all internal relations among the glossing characters. The full network is shown in the following figure.


Eyeballing the network, we can see that the system does look rather systematic. The network is not connected and, apart from a few large connected components, we find a lot of discrete groups that seem to reflect individual initial sounds that were clearly distinguished from other sounds in the fǎnqiè spellings.

The following figure shows a part of the network, namely the second cluster in the big network (above) when going from left to right and staying at the top. In this figure, we can see that the network has two highly connected source characters linking to almost all of the other characters.


I have to admit that I am still having trouble interpreting the network satisfactorily, let alone designing more complex methods to analyse it. Nevertheless, I have the hope that the network analysis of Chinese pronunciation glosses can give us new insights into the phonetic history of Chinese. Importantly, the structures reflected by the network are true pronunciation differences, and that we can indeed find concrete sounds in the indirect fǎnqiè spelling system, becomes specifically clear when comparing the reconstructed pronunciations of the characters in the sample with each other.

For example, when you look at the figure below, you can see that our connected component represents two different clusters of initials, namely a simple k and an aspirated . The node that links the two groups is given the pronunciation in our example, but its original reading is ambiguous. The character has two readings and two meanings reflecting both ancient k and ancient (today pronounced as jiē «Chinese pistachio tree» and kǎi «template», respectively).


Networks of pronunciation glosses in Chinese Traditional Phonology are still under-explored, both with respect to traditional scholarship and with respect to the way they are best handled and analyzed in modern network approaches. If we could develop an approach that would infer the clusters of glosses that point consistently to the same sound, they could give us fascinating insights, not only into the phonological system of Chinese varieties spoken during a given time period, but perhaps also into the dynamics underlying pronunciation changes, when comparing different networks across different times and places.

References
  • Branner, D. (2000) The rime-table system of formal Chinese phonology. In: Auroux, S., E. Koerner, H.-J. Niederehe, and K. Versteegh (eds.): History of the language sciences.1.18. de Gruyter: Berlin and New York. 46-55.
  • Coblin, W. (1983) A Handbook of Eastern Han Sound Glosses. The Chinese University Press: Chicago.
  • Gēng Zhènshēng 耿振生 (2004) 20 shìjì Hànyǔ yǔyīnxué fāngfǎ lùn 20世纪汉语音韵学方法论 [20th century’s methods in traditional Chinese phonology]. Běijīng Dàxué 北京大學: Běijīng 北京.
  • International Phonetic Association (1999): IPA Handbook. Cambridge University Press: Cambridge.
  • Kalusky, W. (2017) Die Transkription der Sprachlaute des Internationalen Phonetischen Alphabets: Vorschläge zu einer Revision der systematischen Darstellung der IPA-Tabelle. LINCOM Europa: München.
  • List, J.-M. (2017) Network approaches to the reconstruction of Old Chinese phonology. Talk, held at the "Center for Chinese Linguistics" (2017/03/07, Hong Kong, The Hong Kong University of Science and Technology).

Wednesday, November 11, 2015

Networks in Chinese poetry


Structure in Poetry

Dealing with poetry is a dangerous topic in science, since we never know whether the structures we propose are really there or not. Once it comes to the search of structure in poetry, Matthew and Luke were right, since the ones who search will find, provided they have enough creativity.

When I had Latin lessons in school, some of my classmates were incredibly diligent in trying to find alliterations (instances in which words in a sentence start with the same letter) in Cicero's speeches. This was less out of interest in the structure of the speeches, but more an attempt to divert the teacher's attention away from translation.

The problem with structure in poetry is that we never know in the end whether the people who created the poetry did things with purpose or not. Consider, for example, the following lines of a famous verse:


Apart from the fact that people might disagree whether songs by Eminem are poetry, it is interesting to look at the structures one may (or may not) detect. We know that rap and hip hop allow for rather loose rhyming schemes, which may give the impression that they were produced in an ad-hoc manner. We know also that the question of what counts as a rhyme is strictly cultural. In German, for example, employ could rhyme with supply (thanks to Goethe and other poets who would superimpose to the standard language rhyme patterns that made sense in their home dialect). If I was given Eminem's poem in an exam, I would mark its rhyming structure as follows:


I do not know whether any teacher of English would agree that music can rhyme with own it, but if Germans can rhyme [ai] (as in supply) with [ɔi] (as in employ), why not allow [ɪk] (as in music) to rhyme with [ɪt] (as in own it)? I bet that if one made an investigation of all rhymes that Bob Dylan has produced so far, we would find at least a few instances where he would tolerate Eminem's rhyme pattern.

The point here is that rhymes are important evidence to infer how Ancient Chinese was pronounced.

The Pronunciation of Ancient Chinese

The Chinese writing system gives only minimal hints regarding the pronunciation of the characters. If one writes a character like 日 which means 'sun', the writing system gives us no clue as to its pronunciation; and from the modern form in which the character is written, it is also difficult to see the image of a sun in the character. Thus, the current situation in Chinese linguistics is that we have very ancient texts, dating at times back to 1000 BC, but we do not have a real clue as to how the language was pronounced by then.

That it was pronounced differently is clear from — ancient Chinese poetry. When reading ancient poems with modern pronunciations, one often finds rhyme patterns which do not sound nice. Consider the poem from Ode 28 of the Book of Odes (Shījīng 詩經), an ancient collection of poems written between 1050 and 600 BC (translation from Karlgren 1950):


Here, we find modern rhymes between fēi and guī which is fine, since the transliteration fails to give the real pronunciation, which is [fəi] versus [kuəi]; but we also find [in] rhyming with [nan], which is so strange (due to the strong difference in the vowels) that even Bob Dylan and Eminem probably would not tolerate it. But if we do not tolerate this rhyming pattern, and if we do not want to assume that the ancient masters of Chinese poetry would simply fail in rhyming, we need to search for some explanation as to why the words do not rhyme. The explanation is, of course, language evolution — The sound systems of languages constantly change, and if things do not rhyme with our modern pronunciation, they may have been perfect rhymes when they were originally created.

When Chinese scholars of the 16th century, who investigated their ancient poetry, became aware of this, they realized that the poetry could be a clue to reconstruct the ancient pronunciation of their language. Then they began to investigate the ancient poems of the Book of Odes systematically for their rhyme patterns. It is thanks to this work on early linguistic reconstruction by Chinese scholars, that we now have a rather clear picture of how Ancient Chinese was pronounced (see especially Baxter 1992, Sagart 1999, and Baxter and Sagart 2014).

Networks in Chinese Rhyme Patterns

But where are the networks in Chinese poetry, which I promised in the title of this post? They are in the rhyme patterns — It is rather straightforward to model rhyme patterns in poetry with the help of networks. Every node is a distinct word that rhymes in at least one poem with another word. Links between nodes are created whenever one word rhymes with another word in a given stanza of a poem. So, even if we take only two stanzas of two poems of the Book of Odes, we can already create a small network of rhyme transitions, as illustrated in the following figure:


One needs, of course, to be careful when modeling this kind of data, since specific kinds of normalizations are needed to avoid exaggerating the weight assigned to specific rhyme connections. It is possible that poets just used a certain rhyme pattern because they found it somewhere else. It is also not yet entirely clear to me how to best normalize those cases in which more than two words rhyme with each other in the same stanza.

But apart from these rather technical questions, it is quite interesting to look at the patterns that evolve from collecting rhyme patterns of all poems found in the Book of Odes, and plotting them in a network. I prepared such a dataset, using the rhyme assessments by Baxter (1992). The whole data set is now available in the form of an interactive web-application at http://digling.org/shijing.

In this application, one can browse all characters that appear in potential rhyme positions in all 305 poems that constitute the Book of Odes. Additional meta-data, like reconstructions for the old pronunciations following Baxter and Sagart (2014), which were kindly provided by L. Sagart, have also been added. The core of the app is the "Poem View", by which one can see a poem, along with reconstructions for the rhyme words, and an explicit account of what experts think rhymed in the classical period, and what they think did not rhyme. The following image gives a screanshot of the second poem of the Book of Odes:



But let's now have a look at the big picture of the network we get when taking all words that rhyme into account. The following image was created with Cytoscape:



As we can see, the rhyme words in the 305 poems almost constitute a small world network, and we have a very large connected component. For me, this was quite surprising, since I was assuming that rhyme patterns would be more distinct. It would be very interesting to see a network of the works of Shakespeare or Goethe, and to compare the amount of connectivity.

There are, of course, many things we can do to analyze this network of Chinese poetry, and I am currently trying to find out to what degree this may contribute to the reconstruction of the pronunciation of Ancient Chinese. But since this work is all in a preliminary stage, I will restrict this post by showing how the big network looks if we color the nodes in six different colors, based on which of the six main vowels ([a, e, i, o, u, ə]) scholars usually reconstruct in the rhyme word for Ancient Chinese:



As can be seen, even this simple annotation shows how interesting structures emerge, and how we see more than before.

Many more things can be done with this kind of data. This is for sure. We could compare the rhyme networks of different poets, maybe even the networks of one and the same poet at different stages of their life, asking questions like: "do people rhyme more sloppy, the older they get?" It's a pity that we don't have the data for this, since we lack automatic approaches to detect rhyme words in text, and there are no manual annotations of poem collections apart from the Book of Odes that I know of.

But maybe, one day, we can use networks to study the dynamics underlying the evolution of literature. We could trace the emergence of rap and hip hop, or the impact of the "Judas!"-call on Dylan's rhyme patterns, or the loss of structure in modern poetry. But that's music from the future, of course.

References
  • Baxter, William H. (1992) A handbook of Old Chinese phonology. Berlin: De Gruyter.
  • Baxter, William H. and Sagart, Laurent (2014) Old Chinese. A new reconstruction. Oxford: Oxford University Press.
  • Karlren, Bernhard (1950) The Book of Odes. Stockholm: Museum of Far Eastern Antiquities.
  • Sagart, Laurent (1999) The roots of Old Chinese. Amsterdam: John Benjamins.