Showing posts with label Language Networks. Show all posts
Showing posts with label Language Networks. Show all posts

Monday, May 28, 2018

Comparing reconstruction systems in historical linguistics


The term linguistic reconstruction has a very specific meaning in historical linguistics, pointing usually to the techniques that are used in order to infer how a given language was originally pronounced, even though it has not been attested in written sources. In previous posts, I have occasionally pointed to reconstructed forms, the so-called proto-forms, which linguists usually mark as such by putting an asterisk in from of them. For example, the word Indo-European *ph₂tér- is a reconstructed proto-form for the supposed Indo-European word "father".

While the reconstruction techniques are usually limited to languages for which we have no written record, they can in principle also be applied in order to find out how ancient languages like, for example, Latin and Greek, were pronounced in detail (Sturtevant 1920). For languages like Chinese, whose writing system leaves almost no clues about pronunciation, linguistic reconstruction is the only way to investigate the pronunciation of the oldest stages of the language.

When dealing with different reconstruction systems for Old Chinese phonology, it is quite difficult, even for experienced scholars, to spot the actual differences between the systems. That these differences exist, and that they can be quite substantial, is beyond question — and easy to understand, if one takes into account that Old Chinese is reconstructed with the help of a philological (as opposed to a mainly comparative) approach, by which data from different sources is sifted and individually weighed (cf. Jarceva 1990: 409 and List 2008).

When comparing different reconstruction systems, it is not enough simply to look at the inventories of proto-phonemes proposed by different scholars. Even if two proto-inventories (the sets of the reconstructed sounds) are exactly the same, it is possible that scholars will provide different reconstructions for individual characters. The only way to compare two or more reconstruction systems is therefore to compare the concrete reconstructions for a certain number of characters.

In addition to the sample of words, however, we also need a clear account of which segments (which proto-sounds) should be compared with each other. When comparing proto-forms for Chinese 一 ‘one’ in different Old Chinese reconstruction systems, such as Karlgren (1950) *ʔi ̯ĕt, Li (1971) *ʔjit, Wáng (1980) *iet, and Baxter and Sagart (2014) *ʔi[t], we would obviously not compare the medial *i ̯ of Karlgren with the initial *ʔ of Baxter and Sagart.

When adding more reconstructions, such as the one for 七 ‘seven’ across the four systems, for which the authors give *ts'i ̯ĕt, *tshjit, *tshiet, and *[tsh]i[t], respectively, we can further see that there are not only differences for the different segments in the same positions, but also for the interpretation of the words. Although all authors give different medials, main vowels, and finals in the two words, they are structurally consistent in giving both words the same sound segments for medial, nucleus, and coda.

What we can see from this example is that any difference in the sound segments, like the choice of initials, or the concrete solution proposed for a problem, does not immediately reflect important differences in the reconstruction systems. If two scholars just choose another symbol for a distinction that they both recognize and acknowledge, this does not render the reconstructions incompatible. It should therefore not be used as a criterion for dismissing a given reconstruction system, at least not in a first step. If two systems are structurally equivalent, then they have equivalent predictive power for the descendant language(s) they are supposed to reconstruct.

This abstractionist notion of proto-forms, which can be found in the early work of Saussure (1916) and Meillet (1903), is problematic for the endeavour of linguistic reconstruction, and usually not strictly followed (Lass 2017). Nevertheless, the potentially abstract notion of proto-forms is important to be kept in mind when comparing different reconstruction systems. When distinguishing the structural differences (which result from the direct interpretation of the data and the identification of regular sound correspondences) from the substantial differences (resulting from a phonetic and phonological interpretation of the identified correspondences), we have a much clearer account of the core of the differences, and whether they are worth our consideration or not.

But how can we compare reconstruction systems structurally? Firstly, we need to have the data assembled in aligned form, in order to make sure that we only compare like with like (e.g., medial with medial, and vowel with vowel). A sample illustration in which alignments of the proto-forms for ‘seven’ and ‘one’, produced with the help of the EDICTOR tool (List 2017), is given in the figure below.

Comparing reconstruction proposals with the help of alignments.

Alternatively, we can also select a single aspect, such as, for example, the vowel system proposed in different reconstruction systems. Having assembled a substantial amount of different proto-forms in this way, the structural comparison between different reconstruction systems can be modeled as a comparison of different cluster analyses, or, more accurately, partitioning analyses. A partitioning analysis assigns a given number of objects to a certain number of different groups. When dealing only with the vowels proposed by different reconstruction systems, we can say that a given reconstruction, like the one by Karlgren, for example, assigns each Chinese character, for which a proto-form is given, to a particular group depending on the main vowel selected for the reconstruction.

If, for a given number of reconstructions, we model each reconstruction system as a partitioning analysis, based on the main vowel proposed by the system, we can use standard metrics from graph theory and Natural Language Processing to compare different reconstruction systems with each other. Very straight-forward measures for the comparison of two partitioning analyses are the so-called B-Cubed scores (Amigó et al. 2009), which have proven specifically useful for the evaluation of automatic cognate detection methods in historical linguistics, compared to a gold standard (Hauer and Kondrak 2011, List et al. 2017).

Being an evaluation measure, B-Cubed scores come in the typical three flavors of precision, recall, and F-Score. Precision is similar to the notion of true positives, and recall is similar to true negatives. For the purpose of comparing reconstruction systems, only the F-score is needed, as it is a symmetric measure, and the notion of true positives and true negatives is meaningless, unless we decide that we blindly trust one of the given systems. As also for the scores for precision and recall, the F-score ranges between 0 and 1, with 1 indicating that the two partitioning analyses are identical.

In order to compare more than one reconstruction system, we can make use of techniques for exploratory data analysis (Morrison 2014); and the most straightforward way to do this, is, of course, to use the NeighborNet algorithm (Bryant and Moulton 2004), as provided by the SplitsTree package (Huson 1998).

In order to illustrate how data-display networks can be used to study differences among Old Chinese reconstruction systems, I designed a little experiment, based on data taken from (List et al. 2017b), who provide Old Chinese reconstructions for all rhyme words in the Shījīng based on eight different reconstruction systems (Baxter and Sagart 2014, Karlgren 1950, Li 1971, Pān 2000, Schuessler 2007, Starostin 1989, Wáng 1980, Zhèngzhāng 2003).

In order to keep the analysis simple, I extracted only the different reconstructions of the main vowel for each character in each system, and carried out a pairwise comparison of all eight systems, computing the B-Cubed F-scores for each pair, omitting characters for which no reconstruction could be found in the data. These scores were then converted to a distance matrix, and fed to the NeighborNet algorithm (the source code can be downloaded here). The resulting network is provided in the figure below.

NeighorNet reflecting the closeness of the different reconstruction systems
As one can see, the data roughly clusters into three subgroups, namely Schuessler, Baxter and Sagart, and Starostin vs. Pān and Zhèngzhāng vs. Karlgren, Li, and Wáng. On a larger scale, we can divide the data into all six-vowel systems versus the non-six-vowel systems (Karlgren, Wáng, Li). Given that Pān is a direct student of Zhèngzhāng, the closeness between their reconstruction systems is not surprising.

What may be surprising is the closeness of the Schuessler, Starostin, and Baxter and Sagart systems, given their notable differences with respect to the criterion of vowel purity tested by List et al. (2017b). Even if the network analysis cannot directly explain all of these differences in detail, it seems like a worthwhile enterprise, which should be further expanded by comparing not only the vowels, but fully aligned proto-forms.

Given the straightforwardness of the application, it seems also useful to test it on other language families where there is similar disagreement, as in the reconstruction of Old Chinese phonology.

References

Amigó, E., J. Gonzalo, J. Artiles, and F. Verdejo (2009): A comparison of extrinsic clustering evaluation metrics based on formal constraints. Information Retrieval 12.4. 461-486.

Baxter, W. and L. Sagart (2014) Old Chinese: a new reconstruction. Oxford University Press: Oxford.

Bryant, D. and V. Moulton (2004) Neighbor-Net. An agglomerative method for the construction of phylogenetic networks. Molecular Biology and Evolution 21.2. 255-265.

Hauer, B. and G. Kondrak (2011) Clustering semantically equivalent words into cognate sets in multilingual lists. In: Proceedings of the 5th International Joint Conference on Natural Language Processing. AFNLP 865-873.

Huson, D. (1998) SplitsTree: analyzing and visualizing evolutionary data. Bioinformatics 14.1. 68-73.

Jarceva, V. (1990) Sovetskaja Enciklopedija: Moscow.

Karlgren, B. (1950) The Book of Odes. Chinese text, transcription and translation. Museum of Far Eastern Antiquities: Stockholm.

Lass, R. (2017) Reality in a soft science: the metaphonology of historical reconstruction. Papers in Historical Phonology 2.1. 152-163.

Li Fang-kuei 李方桂 (1971) Shànggǔyīn yánjiū 上古音研究 [Studies on Archaic Chinese phonology]. Qīnghuá Xuébào 清華學報 9.1-2. 1-60.

List, J.-M. (2008) Rekonstruktion der Aussprache des Mittel- und Altchinesischen. Vergleich der Rekonstruktionsmethoden der indogermanischen und der chinesischen Sprachwissenschaft [Reconstruction of the pronunciation of Middle and Old Chinese. Comparison of reconstruction methods in Indo-European and Chinese linguistics]. Magister thesis. Freie Universität Berlin: Berlin.

List, J.-M., S. Greenhill, and R. Gray (2017) The potential of automatic word comparison for historical linguistics. PLOS One 12.1. 1-18.

List, J.-M. (2017) A web-based interactive tool for creating, inspecting, editing, and publishing etymological datasets. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics. System Demonstrations. 9-12.

List, J.-M., J. Pathmanathan, N. Hill, E. Bapteste, and P. Lopez (2017) Vowel purity and rhyme evidence in Old Chinese reconstruction. Lingua Sinica 3.1. 1-17.

Meillet, A. (1903) Introduction à l’étude comparative des langues indo-européennes. Hachette: Paris.

Morrison, D.A. (2014) Phylogenetic networks: a new form of multivariate data summary for data mining and exploratory data analysis. WIREs Data Mining and Knowledge Discovery 4: 296-312.

Pān Wùyún 潘悟云 (2000) Hànyǔ lìshǐ yīnyùnxué 汉语历史音韵学 [Chinese historical phonology]. Shànghǎi Jiàoyù 上海教育: Shànghǎi 上海.

de Saussure, F. (1916) Cours de linguistique générale. Payot: Lausanne.

Schuessler, A. (2007) ABC Etymological dictionary of Old Chinese. University of Hawai’i Press: Honolulu.

Starostin, S. (1989) Sravnitel’no-istoričeskoe jazykoznanie i leksikostatistika [Comparative-historical linguistics and lexicostatistics]. In: Kullanda, S., J. Longinov, A. Militarev, E. Nosenko, and V. Shnirel’man (eds.): Lingvističeskaja rekonstrukcija i drevnejšaja istorija VostokaMaterialy k diskussijam na konferencii.[Materials for the discussion on the conference].1. Institut Vostokovedenija: Moscow. 3-39.

Sturtevant, E. (1920) The pronunciation of Greek and Latin. University of Chicago Press: Chicago.

Zhèngzhāng Shàngfāng 郑张尚芳 (2003) Shànggǔ yīnxì 上古音系 [Old Chinese phonology]. Shànghǎi Jiàoyù 上海教育: Shànghǎi 上海.

Tuesday, December 20, 2016

Isogloss maps are hypergraphs are bipartite networks


Linguists are a very special people. They are very proud, especially when biologists tell them how to do phylogenetic analyses; but their pride is often also justified, as many phylogenetic concepts were initially or independently developed by linguists, be it the family tree model, proposed years before Darwin's (1859) tree by Ćelakovský (1853), or even the cladistic principle of synapomorphies, which are called "exclusively shared innovations" in linguistics (see Brugmann 1884).

Linguists also invented one interesting kind of data-display which so far has never been used by biologists (at least as far as I know): maps of isogloss boundaries. The term "isogloss" is an unfortunate term, as it has multiple usages in linguistics, and its history seems to go back to a naive borrowing from chemistry (but I have not really followed the literature here). On most occasions, it just means "shared trait". That is, it denotes a features shared between two or more languages; and given that languages may share many different features, isoglosses for a group of related languages may yield a very complex type of data. Isoglosses are somehow related to the wave theory, the arch-enemy of the family tree in linguistics, which I described as a mystical theory some time ago, since it never really made it to a clear-cut model that could be formalized (The Wave Theory: the predecessor of network thinking in historical linguistics ).

Some linguists, nevertheless, insist that the waves that are the core of the wave theory are nothing other than isoglosses. More specifically, the waves represent innovations that contribute to the separation of languages (a change in pronunciation of a word here, a change in grammar there), but which are not transmitted vertically — they spread across the speakers of a language and may even cross linguistic borders. One early visualization of these waves can be found in Bloomfield (1933), as shown here:


What Bloomfield essentially does here is pick certain traits of Indo-European languages, calling them isoglosses, and arrange them on a quasi-geographic map of Indo-European languages in such a way that all languages sharing a trait are inside one of these isogloss boundaries.

Only recently, I realised, what this actually means, when I found the "Bible of Network Theory" by Newman (2010) and started reading at a random page, which — as it turned out — treated hypergraphs. Hypergraphs, as I learned from Newman, are graphs in which one edge can connect to more than one node, and Newman used exactly the same visualization for these hyperedges as Bloomfield had done in 1933, without knowing that it was actually a rather complex network structure he was proposing.

Even more interesting than the complex graph structure is that hypergraphs can be likewise displayed as bipartite networks, in which we distinguish two fundamental kinds of nodes, and in which connections are only allowed between nodes of different kinds, without losing any information. In order to do so, one just converts all hyperedges into a node that connects to all nodes (languages in our case) to which the edges connect in the hypergraph. In the same way that Bloomfield labeled the hyperedges in his legend, we can label the isogloss nodes that connect to the languages. The following image shows the resulting bipartite network for Bloomfield's hypergraph:


If you now ask what this tells us after all, I will disappoint you — so far it does not tell us anything, it is just a display of data in a different fashion. Note, however, that hypergraph visualization is not a trivial problem, and if you have enclaves not sharing a trait, it may even be impossible to visualize hypergraphs in a two-dimensional space by just using one line that connects to all nodes. Bipartite networks are easier to handle in this regard. Even more importantly, however, bipartite graphs are also easy to handle algorithmically, and biologists are currently developing new methods to handle them (Corel et al. 2016).

If we visualize the Bloomfield data in a bipartite network using network visualization software such as Cytoscape, we can conveniently explore the data, and arrange the nodes in order to search for patterns in the isoglosses. The following visualization, for example, shows that Bloomfield chose the data well in order to illustrate the amount of conflicting, apparently non-tree-like, signal in Indo-European languages (remember that linguists tend to dislike trees, but not necessarily in a productive way), as the data describes more of a circular structure than a strict hierarchy.


In order to really interpret this kind of data, however, we should not forget that this is still a data-display network. It is by no means a phylogenetic analysis, as we only show how a certain amount of data selected by a scholar and distributed over the given language groups. A true phylogenetic analysis will need to interpret these data, making bold claims about the history of those shared traits.

The existence of sibilants (s-like sounds, like [s, z, ʃˌ ʒ]) for certain velar sounds (k-like sounds, like [k, g, x]), for example, is a trait shared by Balto-Slavic, Indo-Iranian, Armenian, and Albanian, but this does not mean that they all inherited it from a common ancestor, as the process of palatalization, by which velar sounds turn into affricates and fricatives (compare French cent, which was pronounced kentum in Latin), is very frequent in the languages of the world, and may well reflect independent evolution.

Apart from independent development, which would actually force us to revise our network, deleting the respective edges because they are not homologous in the strict sense means that we may also have to deal with differential loss. This quite likely happened with the shared feature labeled as "past e-" in the network, referring to the past tense in Ancient Greek and Indo-Iranian, which was augmented by the prefix e-.

A further reason for those commonalities labelled as isoglosses by linguists may also be simple lateral transfer due to language contact.

Proponents of the wave theory have taken this kind of data as proof that the family tree model is essentially wrong. While I would agree that the family tree model shows only a certain aspect of language evolution, and may therefore be boring at times (and even wrong, if we do not manage to correctly interpret the nature of shared traits), I have a hard time understanding why linguists still insist that isogloss maps are an alternative model of language evolution. They are surely not, in the same way in which splits graphs are not phylogenetic networks, as David emphasized in a recent blogpost.

Unless we add the missing time dimension and analyse how the shared traits originated, isogloss maps and hypergraphs will remain nothing more than an interesting form of data visualization. Given the recent research on bipartite networks, however, we may have some hope that the mysterious waves in historical linguistics may not only find a formal model of representation, but even bring us to the point where we gain new insights into the history of our languages.

References
  • Bloomfield, L. (1973) Language. Allen & Unwin: London.
  • Brugmann, K. (1884) Zur Frage nach den Verwandtschaftsverhältnissen der indogermanischen Sprachen [Questions regarding the closer relationship of the Indo-European languages]. Internationale Zeischrift für allgemeine Sprachewissenschaft 1. 228-256.
  • Čelakovský, F. (1853) Čtení o srovnavací mluvnici slovanské [Lectures on comparative grammar of Slavic]. V komisí u F. Řivnáče: Prague.
  • Corel, E., P. Lopez, R. Méheust, and E. Bapteste (2016) Network-thinking: graphs to analyze microbial complexity and evolution. Trends Microbiol. 24.3: 224-237.
  • Darwin, C. (1859) On the origin of species by means of natural selection, or, the preservation of favoured races in the struggle for life. John Murray: London.
  • Newman, M. (2010) Networks. An Introduction. Oxford University Press: Oxford.

Wednesday, November 11, 2015

Networks in Chinese poetry


Structure in Poetry

Dealing with poetry is a dangerous topic in science, since we never know whether the structures we propose are really there or not. Once it comes to the search of structure in poetry, Matthew and Luke were right, since the ones who search will find, provided they have enough creativity.

When I had Latin lessons in school, some of my classmates were incredibly diligent in trying to find alliterations (instances in which words in a sentence start with the same letter) in Cicero's speeches. This was less out of interest in the structure of the speeches, but more an attempt to divert the teacher's attention away from translation.

The problem with structure in poetry is that we never know in the end whether the people who created the poetry did things with purpose or not. Consider, for example, the following lines of a famous verse:


Apart from the fact that people might disagree whether songs by Eminem are poetry, it is interesting to look at the structures one may (or may not) detect. We know that rap and hip hop allow for rather loose rhyming schemes, which may give the impression that they were produced in an ad-hoc manner. We know also that the question of what counts as a rhyme is strictly cultural. In German, for example, employ could rhyme with supply (thanks to Goethe and other poets who would superimpose to the standard language rhyme patterns that made sense in their home dialect). If I was given Eminem's poem in an exam, I would mark its rhyming structure as follows:


I do not know whether any teacher of English would agree that music can rhyme with own it, but if Germans can rhyme [ai] (as in supply) with [ɔi] (as in employ), why not allow [ɪk] (as in music) to rhyme with [ɪt] (as in own it)? I bet that if one made an investigation of all rhymes that Bob Dylan has produced so far, we would find at least a few instances where he would tolerate Eminem's rhyme pattern.

The point here is that rhymes are important evidence to infer how Ancient Chinese was pronounced.

The Pronunciation of Ancient Chinese

The Chinese writing system gives only minimal hints regarding the pronunciation of the characters. If one writes a character like 日 which means 'sun', the writing system gives us no clue as to its pronunciation; and from the modern form in which the character is written, it is also difficult to see the image of a sun in the character. Thus, the current situation in Chinese linguistics is that we have very ancient texts, dating at times back to 1000 BC, but we do not have a real clue as to how the language was pronounced by then.

That it was pronounced differently is clear from — ancient Chinese poetry. When reading ancient poems with modern pronunciations, one often finds rhyme patterns which do not sound nice. Consider the poem from Ode 28 of the Book of Odes (Shījīng 詩經), an ancient collection of poems written between 1050 and 600 BC (translation from Karlgren 1950):


Here, we find modern rhymes between fēi and guī which is fine, since the transliteration fails to give the real pronunciation, which is [fəi] versus [kuəi]; but we also find [in] rhyming with [nan], which is so strange (due to the strong difference in the vowels) that even Bob Dylan and Eminem probably would not tolerate it. But if we do not tolerate this rhyming pattern, and if we do not want to assume that the ancient masters of Chinese poetry would simply fail in rhyming, we need to search for some explanation as to why the words do not rhyme. The explanation is, of course, language evolution — The sound systems of languages constantly change, and if things do not rhyme with our modern pronunciation, they may have been perfect rhymes when they were originally created.

When Chinese scholars of the 16th century, who investigated their ancient poetry, became aware of this, they realized that the poetry could be a clue to reconstruct the ancient pronunciation of their language. Then they began to investigate the ancient poems of the Book of Odes systematically for their rhyme patterns. It is thanks to this work on early linguistic reconstruction by Chinese scholars, that we now have a rather clear picture of how Ancient Chinese was pronounced (see especially Baxter 1992, Sagart 1999, and Baxter and Sagart 2014).

Networks in Chinese Rhyme Patterns

But where are the networks in Chinese poetry, which I promised in the title of this post? They are in the rhyme patterns — It is rather straightforward to model rhyme patterns in poetry with the help of networks. Every node is a distinct word that rhymes in at least one poem with another word. Links between nodes are created whenever one word rhymes with another word in a given stanza of a poem. So, even if we take only two stanzas of two poems of the Book of Odes, we can already create a small network of rhyme transitions, as illustrated in the following figure:


One needs, of course, to be careful when modeling this kind of data, since specific kinds of normalizations are needed to avoid exaggerating the weight assigned to specific rhyme connections. It is possible that poets just used a certain rhyme pattern because they found it somewhere else. It is also not yet entirely clear to me how to best normalize those cases in which more than two words rhyme with each other in the same stanza.

But apart from these rather technical questions, it is quite interesting to look at the patterns that evolve from collecting rhyme patterns of all poems found in the Book of Odes, and plotting them in a network. I prepared such a dataset, using the rhyme assessments by Baxter (1992). The whole data set is now available in the form of an interactive web-application at http://digling.org/shijing.

In this application, one can browse all characters that appear in potential rhyme positions in all 305 poems that constitute the Book of Odes. Additional meta-data, like reconstructions for the old pronunciations following Baxter and Sagart (2014), which were kindly provided by L. Sagart, have also been added. The core of the app is the "Poem View", by which one can see a poem, along with reconstructions for the rhyme words, and an explicit account of what experts think rhymed in the classical period, and what they think did not rhyme. The following image gives a screanshot of the second poem of the Book of Odes:



But let's now have a look at the big picture of the network we get when taking all words that rhyme into account. The following image was created with Cytoscape:



As we can see, the rhyme words in the 305 poems almost constitute a small world network, and we have a very large connected component. For me, this was quite surprising, since I was assuming that rhyme patterns would be more distinct. It would be very interesting to see a network of the works of Shakespeare or Goethe, and to compare the amount of connectivity.

There are, of course, many things we can do to analyze this network of Chinese poetry, and I am currently trying to find out to what degree this may contribute to the reconstruction of the pronunciation of Ancient Chinese. But since this work is all in a preliminary stage, I will restrict this post by showing how the big network looks if we color the nodes in six different colors, based on which of the six main vowels ([a, e, i, o, u, ə]) scholars usually reconstruct in the rhyme word for Ancient Chinese:



As can be seen, even this simple annotation shows how interesting structures emerge, and how we see more than before.

Many more things can be done with this kind of data. This is for sure. We could compare the rhyme networks of different poets, maybe even the networks of one and the same poet at different stages of their life, asking questions like: "do people rhyme more sloppy, the older they get?" It's a pity that we don't have the data for this, since we lack automatic approaches to detect rhyme words in text, and there are no manual annotations of poem collections apart from the Book of Odes that I know of.

But maybe, one day, we can use networks to study the dynamics underlying the evolution of literature. We could trace the emergence of rap and hip hop, or the impact of the "Judas!"-call on Dylan's rhyme patterns, or the loss of structure in modern poetry. But that's music from the future, of course.

References
  • Baxter, William H. (1992) A handbook of Old Chinese phonology. Berlin: De Gruyter.
  • Baxter, William H. and Sagart, Laurent (2014) Old Chinese. A new reconstruction. Oxford: Oxford University Press.
  • Karlren, Bernhard (1950) The Book of Odes. Stockholm: Museum of Far Eastern Antiquities.
  • Sagart, Laurent (1999) The roots of Old Chinese. Amsterdam: John Benjamins.