Chinese lexical database
Cluster 16: character 2 homographs
C2 Homographs (Types) 0.29 0 0.55 0 3 4710
C2 Homographs (Tokens) 4.20 0 21.93 0 366 4710
C2 Homographs Freq. 76.74 0.00 694.78 0.00 45938.18 4710 Cluster 17: character 1 phonetic radical enemies
C1 PR Enemies (Types) 3.43 3 2.91 0 14 14478
C1 PR Enemies (Tokens) 47.12 26 55.56 0 302 14478
C1 PR Enemies Freq 767.93 150.12 2143.75 0.00 46504.14 14478
C1 PR Family Size 6.70 6 4.57 1 19 14478
C1 PR Frequency 976.60 449.70 1612.30 0.00 28622.72 14478 Cluster 18: character 2 phonetic radical enemies
C2 PR Enemies (Types) 3.44 3 2.88 0 14 17944
C2 PR Enemies (Tokens) 46.60 25 54.78 0 302 17944
C2 PR Enemies Freq 886.95 153.82 2951.46 0.00 46501.34 17944
C2 PR Family Size 6.61 6 4.55 1 19 17944
Character 1 Homographs Tokens is the number of words in which the first character is pronounced differently. For the first character差 in the word 差使, the alternative pronunciations [thù¨a1], [thù¨a4], and [thsi1] occur in 17, 4 and 1 words, respectively. Therefore, Character 1 Homographs (Tokens) for the word 差使 is 17 + 4 + 1 = 22.
Finally, Character 1 Homographs Frequency is the summed frequency of all character 1 homograph tokens. The summed frequency of the 22 homograph tokens for the character差 in the word 差使 is 247.56. Character 1 Homograph Frequency for the word差使, therefore, is 247.56.
Table 2.22 presents the pairwise correlations between the measures in Cluster 15.
All correlations are near-perfect and highly significant at the 0.001 α level. The 3 measures Character 1 Homographs (Types), Character 1 Homographs (Tokens), Character 1 Homographs Frequency thus encode very similar information.
2.4. NUMERICAL VARIABLES 61
Table 2.22: Pairwise Spearman correlations for the numerical variables in Cluster 15.
Abbreviations: C1HTY = Character 1 Homographs (Types), C1HTO = Character 1 Homographs (Tokens), C1HF = Character 1 Homograph Frequency.
predictor C1HTY C1HTO C1HF
C1HTY
-C1HTO 0.993
-C1HF 0.991 0.995
-2.4.5.2 Cluster 16: character 2 homographs
Cluster 16 is the character 2 counterpart of Cluster 15 and contains 3 measures related to the number and frequency of homographs for character 2 (Character 2 Homographs (Types), Character 2 Homographs Tokens and Character 2 Homo-graphs Frequency). Table 2.23 presents the pairwise Spearman correlations for the measures in Cluster 16. Like the measures in Cluster 15, the numerical vari-ables in Cluster 16 show near-perfect correlations that are highly significant at an α level of 0.001. Similar to the measures in Cluster 15, therefore, the variables in Cluster 16 encode highly similar information. The pairwise correlations for Cluster 15 and Cluster 16 correlate strongly (r = 0.985). The distributional spaces for the homography measures, thus, are highly similar for both characters.
Yet, there is a subtle but important difference between the homography mea-sures for the first and the second character. The distribution of homography across characters seems fine-tuned to the information-theoretic properties of the immediate linguistic context in which a character appears. The means for all three character 2 measures (Character 2 Homographs (Types): 0.29, Character 2 Homographs (Tokens): 4.20, Character 2 Homographs Frequency: 76.74) are higher than the corresponding means for the character 1 measures for two-character words in Cluster
Table 2.23: Pairwise Spearman correlations for the numerical variables in Cluster 16.
Abbreviations: C2HTY = Character 2 Homographs (Types), C2HTO = Character 2 Homographs (Tokens), C2HF = Character 2 Homograph Frequency.
predictor C2HTY C2HTO C2HF
C2HTY
-C2HTO 0.985
-C2HF 0.983 0.993
-15 (Character 1 Homographs (Types): 0.25, Character 1 Homographs (Tokens):
1.99, Character 1 Homographs Frequency: 51.25). As indicated by paired t-tests for the first and second character homography measures for all two-character words, these differences are significant (types: t(25934) = -9.89, p < 0.001; tokens: t(25934)
= -14.47, p < 0.001; frequency: t(25934) = -4.74, p < 0.001). This suggests that the information provided by the first character reduces the uncertainty about the identity of the second character to such an extent that more variation is possible for the pronunciation of the second character.
Furthermore, characters that form single-character words (Character 1 Ho-mographs (Types): 0.11, Character 1 HoHo-mographs (Tokens): 0.69, Character 1 Homographs Frequency: 10.86) show less homography than first characters in two-character words (Character 1 Homographs (Types): 0.25, Character 1 Ho-mographs (Tokens): 1.99, Character 1 HoHo-mographs Frequency: 51.25). Again, these differences are significant (types: t(8805.69) = -22.67, p < 0.001; tokens:
t(13322.63) = -12.18, p < 0.001; frequency: t(28174.99) = -10.85, p < 0.001). In the context of the information provided by a second character, therefore, the first character is allowed to provide less conclusive information about its pronunciation than when it appears by itself.
These observations suggest that when the uncertainty is sufficiently reduced, the phonological form of a character is allowed to vary. When it is not, a character is preferred to map onto a single phonological form. This fits well with discrimination learning approaches, in which the distributional properties of the language process-ing system are shaped by the need to reduce uncertainty about the lprocess-inguistic input (see, e.g., Ramscar et al., 2013).
2.4.5.3 Cluster 17: character 1 phonetic radical orthography-to-phonology consistency
Clusters 15 and 16 contain numerical variables regarding the orthography-to-phonology consistency at the character level. The measures in Clusters 17 (charac-ter 1) and Clus(charac-ter 18 (charac(charac-ter 2) encode information about the orthography-to-phonology consistency of the phonetic radical. Cluster 17 consists of 5 measures.
The first 3 measures are the counterparts of the measures in Cluster 15 at the pho-netic radical level: Character 1 PR Enemies (Types), Character 1 PR Enemies (Tokens), and Character 1 PR Enemies Frequency.
2.4. NUMERICAL VARIABLES 63 Character 1 PR Enemies (Types) is the number of different pronunciations of characters in which the phonetic radical of the first character appears. For ex-ample, the phonetic radical of the first character 端 in the word 端倪 (“clue”,
“[tu¨an1ni2]”) is 耑. In addition to “[tu¨an1]”, there are 4 other pronunciations of characters that contain this phonetic radical: “[thùuaI4]” (e.g., in 踹, “to kick”,
“[thùuaI4]”), “[thùuAn3]” (e.g., in the first character 喘 of the word 喘息, “to pant”,
“[thùuAn3thCi4]”), “[thuAn1]” (e.g., in the first character湍 of the word 湍流, “rush-ing water”, “[thuAn1liu2]”), and “[üui4]” (e.g., in瑞, “propitious”, “[üui4]”). There-fore, Character 1 PR Enemies (Types) for the word 端倪 is 4.
Character 1 PR Enemies (Tokens) refers to the number of words in which the character that has the same phonetic radical as the first character of the current word is pronounced differently than the first character in the current word. The phonetic radical耑is pronounced as “[thùuaI4]” in 1 word, as “[thùuAn3]” in 5 words, as “[thuAn1]” in 3 words and as “[üui4]” in 2 words. Character 1 PR Enemies (Tokens) for the word 端倪, therefore, is 1 + 5 + 3 + 2 = 11.
Character 1 PR Enemies Frequency is the summed frequency of the enemy tokens. The summed frequency of the 11 words that are phonetic radical enemies of the first character in the word端倪 is 20.37. Character 1 PR Enemies Frequency, therefore, is 20.37.
Compared to the character-level orthography-to-consistency measures (Character 1 Homographs (Types): 0.23, Character 1 Homographs (Tokens): 1.79, Charac-ter 1 Homographs Frequency: 45.04), the corresponding phonetic radical measures have higher means (Character 1 PR Enemies (Types): 3.43, Character 1 PR Ene-mies (Tokens): 47.12, Character 1 PR EneEne-mies Frequency: 767.93). Homography at the phonetic radical level, therefore, is much more common than at the character level.
The fourth measure in Cluster 17 is Character 1 PR Family Size, which is defined as the number of characters the phonetic radical of the first character occurs in. For the word端倪, for instance, the phonetic radical耑 of the first character 端, appears in 5 characters (端, 踹, 喘, 湍, 瑞). Character 1 PR Family Size for the word端倪 thus is 5. The final numerical variable in Cluster 17 is Character 1 PR Frequency, which is the first character equivalent of Character 2 PR Frequency (see the discussion of Cluster 6).
Table 2.24: Pairwise Spearman correlations for the numerical variables in Cluster 17. Abbreviations: C1PRENTY = Character 1 PR Enemies (Types), C1PRENTO
= Character 1 PR Enemies (Tokens), C1PRENFR = Character 1 PR Enemies Frequency, C1PRFS = Character 1 PR Family Size, C1PRF = Character 1 PR Frequency.
predictor C1PRENTY C1PRENTO C1PRENFR C1PRFS C1PRF
C1PRENTY
-C1PRENTO 0.844
-C1PRENFR 0.752 0.924
-C1PRFS 0.847 0.795 0.684
-C1PRF 0.588 0.726 0.783 0.543
-Table 2.24 presents the pairwise correlations for the variables in Cluster 17. All pairwise correlations are positive and significant at an α level of 0.001. The lowest pairwise correlation is the correlation between Character 1 PR Frequency and Character 1 PR Family Size. This correlation is no less than 0.543. Cluster 17, therefore, is a highly homogeneous cluster.
2.4.5.4 Cluster 18: character 2 phonetic radical orthography-to-phonology consistency
Cluster 18 is the character 2 counterpart of Cluster 17 and contains the numerical variables Character 2 PR Enemies (Types), Character 2 PR Enemies (Tokens), Character 2 PR Enemies Frequency, and Character 2 PR Family Size. The pair-wise correlations for the numerical variables in Cluster 18 are presented in Table 2.25.
All correlations are positive and significant at an α level of 0.001. Furthermore, the
Table 2.25: Pairwise Spearman correlations for the numerical variables in Cluster 18. Abbreviations: C2PRENTY = Character 2 PR Enemies (Types), C2PRENTO
= Character 2 PR Enemies (Tokens), C2PRENFR = Character 2 PR Enemies Frequency, C2PRFS = Character 2 PR Family Size.
predictor C2PRENTY C2PRENTO C2PRENFR C2PRFS
C2PRENTY
-C2PRENTO 0.842
-C2PRENFR 0.749 0.913
-C2PRFS 0.837 0.793 0.671
-2.4. NUMERICAL VARIABLES 65 pairwise correlations for Cluster 18 are highly similar to the pairwise correlations for the corresponding measures in Cluster 17 (r = 0.998). This indicates that the distributional structure of the phonetic radical orthography-to-consistency measures is similar for character 1 and character 2.
2.4.6 Group 5: homophones
The clusters in Group 4 contained measures describing the orthography-to-phonology consistency at the level of the character and the phonetic radical. The clusters in Group 5 describe consistency in the other direction: from phonology to orthography. As can be seen in Table 2.26, Cluster 19 describes the phonology-to-orthography consistency for character 1, both at the character-level and at the level of the phonetic radicals. Cluster 20 contains phonology-to-orthography measures for character 2.
2.4.6.1 Cluster 19: character 1 homophones
Cluster 19 consists of 6 phonology-to-orthography consistency measures for char-acter 1. The first 3 measures encode information about the number of homophones
Table 2.26: Overview of numerical predictors: homophones (Group 5)
mean median sd min max NA