• No results found

Comparison of internal consistency

Cross-Entropy

7.2 Comparison of internal consistency

In this section, I compare the internal consistency and precision of the three approaches. If a metric is accurate, it will consistently produce the same distance when presented with a given

language pair. If it is precise, it will consistently rank two pairs of languages which have very similar distances.

The parameter-based approach relies on a single set of values for each language, so it is not inherently variable. However, inconsistencies may arise since establishing those values is sub-ject to researcher fallibility. Firstly, if a lexicon is unrepresentative of the language it is drawn from, other lexicons may produce different parameter values. This can be mitigated by the in-clusion of frequency data, but this is not available for the under-documented languages which are most likely to have short and potentially unrepresentative lexicons, and for which errors are least likely to be caught by peer-review. Secondly, marginal items may be treated inconsistently between languages, being permitted to influence a parameter-value in some cases and not in others. Finally, a user who has specialist knowledge of particular phenomena in one language but not another may selectively deviate from diagnostic criteria. Nidaba contains several tools to mitigate the influence of user variability by automating certain processes, but relying on these to the exclusion of expert knowledge would remove an important verification step.

The resolution of the parameter-based metric is dependent on the number of parameters applicable to a given language pair. The language pair with the smallest number in my sample had 41 applicable parameters, so the metric has a precision of 0.025, and can distinguish between 41 distances. Since no language pairs in my sample are antithetical - something that would be highly unlikely to occur by chance even including thousands of languages - the range of distances observed is 0.06-0.40. This corresponds to approximately 13 distinct categories of language dis-tance. Increasing the number of parameters would increase the precision of the metric.

The entropy-based approach requires transcribed texts to act as exemplars of the language;

one to train a model, and one to test against. The accuracy of the metric therefore depends on how representative these texts are of the language as a whole. The results presented here used translations of a single text for all languages to eliminate confounds such as author- or genre-based variations in entropy. In future, it would be good to repeat the calculations using a variety of source texts, to examine the impact this has on entropy-based language distance metrics.

The results were cross-validated, by repeating the same calculation of Kullback-Leibler di-vergence on multiple sample texts. For all four representational approaches examined, the variation observed between repetitions had a magnitude below 13% of the range of language

distances calculated (see Subsection 5.6.9). Unlike the parameter-based approach, it is there-fore not possible to consistently rank up to 41 distinct language distances (which would require a precision of±1.25%), nor ever the 21 language pairs used in the entropy calculations (re-quiring < ±2.5%). Instead, it is possible to consistently divide language pairs into five non-overlapping groups using the entropy approach, regardless of which of the four transcription methods is used. With only seven languages under examination, it is quite possible that there exist language pairs with greater, or even lesser, language distance between them than we have seen here. In that case, the number of non-overlapping groups would increase. However, since Kullback-Leilber divergence has a fixed normalisation, extending the observed values for the metric would not alter the existing values, and the 21 language pairs examined here will never have fully distinguishable distances using this metric with the transcription systems described.

It is possible that the precision and reliability of the metric could be improved with different representational choices, or with more advanced entropic calculations.

The ACCDIST approach does not have high internal consistency. As with the entropy-based metrics, altering the source data for a language can alter the resulting language distance. How-ever, the entropy-based metric successfully established a minimum data requirement, above which a language could be reliably identified. This is not the case for the ACCDIST approach, where five of the 21 English speakers were more similar to Greek speakers than to their colin-guals.

The ACCDIST metric has a resolution of only three statistically distinct language distance categories: ‘colingual’, ‘similar’, and ‘dissimilar’. Looking at the six non-colingual language pairs that all three approaches have in common, this is the same resolution as three of the four entropy-based metrics. However, these all include German-English in the ‘similar’ category along with Greek-Spanish, which ACCDIST does not (see Table 7.1). By contrast, the entropy metric depend-ing on language-specific binary features divides the language pairs not into two, but into three categories: Greek-Spanish is the closest, followed by Greek-English, with German-English hav-ing a comparable distance to German-Greek or Spanish-English. Finally, the parameter-based approach sorts all six language pairs into distinct categories: German-English is closest, followed by Greek-German, then Greek-Spanish, Greek-English, Spanish-English, and Spanish-German.

Table 7.1: Categorisation of English (Eng.), German, Greek and Spanish (Spa.) by different metrics

Entropy: Entropy: Entropy: Entropy:

ACCDIST IPA static language-specific Elements Parameters

Greek-Spa. Greek-Spa.

Moving on from internal consistency, we can now ask: how similar are the results of the different metrics to each other?

Table 7.2 shows the Pearson correlation between all six metrics. ACCDIST is included in the table for completeness, but only has six data points to the others’ 21, and has been discussed above.

Figure 7.1 comprises six heatmaps showing the relative similarity between languages pro-duced by the parametric Hamming distance, by the mean Kullback-Leibler divergence of each of the four different representational approaches, and by ACCDIST. It includes the 21 language pairs for which the Kullback-Leibler calculations were performed.

The strongest correlation is, unsurprisingly, between the IPA-representation entropy-based metric and the static binary features-representation entropy-based metric. The binary features map directly onto the IPA, and entropy was calculated from abstract segments which therefore closely correspond between the two.