We have argued that sister-head parsing is more useful than head-head parsing because of the NEG RA annotation style. It would be insightful to test how well the parser works with less flat annotations, but on the same data. This would also provide is with clues to another question: why is the increase in performance due to lexicalization much greater in WS J parsing than we have found in NEG RA parsing?
S trictly speaking, it would not be possible to do this without manually re- annotating the corpus. We can, however, approximate a less flat annotation scheme by semi-automatically modifying the corpus so that the annotations more closely resemble those of the WS J. D oing this allows us to test if it really is the annotation style that causes the difference between the head-head and sister-head parser.
This re-annotation has another purpose, as well. G iven that we use the same evaluation metrics that have become common in WS J parsing, it is natural to want to compare our results on NEG RA with known results on the WS J. But are the numbers really comparable? From our attempt to unflatten P Ps in S ection 3. 3, we have some evidence this is not the case. B ut how great is the impact of dependency-like annotation on the evaluation metrics? In this section, we investi- gate this question, as well as the effect of dependency-style annotation on the two lexicalization strategies.
3 . 5 . 1 M etho d
We have already investigated the impact of using some WS J-style annotation, i. e. by applying Rule ( 3. 4) to unflatten PP s. We propose three additional tree trans- formations, all affecting S categories. First, we introduce a VP category bounding the finite verb:
S NP-S B C1 Cn S NP-S B VP C1 Cn ( 3. 6) 47 Lexicalized Parsing
Although the parser we use in this section does not handle traces, one ought be inserted when the sub ject does not occupy position i ( see S ection 4. 4 for more about topicalization) .
The second transformation is to add an S BAR layer in complementizer phrases: S KO US C1 Cn S B AR KO US S C1 Cn ( 3. 7)
We treat subordinating co-ordinators ( KO US ) and relative pronouns ( P RELS , P RELAT and P WAV) as complementizers. Normally, the presence of a comple- mentizer is both a sufficient and necessary condition for an S BAR, but there is one exception: co-ordination. For example, consider the sentence:
Wir reden [S BAR weil Ich du mm b in ] u nd [S BAR du verru ckt¨ bist ]
We talk [ because I stupid am ] and [ you crazy are ]
The complementizer is an empty element in the second S BAR. It can only be detected because it is a co-ordinate sister of the first one.
The third and final change involves pronouns and nouns. NEG RA does not contain unary productions, so any pronouns and singleton nouns will attach directly to an S node ( or VP node) without the benefit of an intermediary NP node. We re-introduce these nodes. An example of the last transformation would be: S PP ER C0 Cn S NP P PER C1 Cn ( 3. 8 )
Note, though, that in addition to PP ER tags, we also invoke this operation when any pronoun or stand-alone noun tag are found.
Each of the tree transformations aboves, along with the P P transformation from S ection 3. 3 are applied one at a time to the sister-head parser in the perfect tags condition.
P recision Recall F-score Avg C B 0C B 6 2BC Cov Baseline 73. 8 74. 4 74. 1 0. 65 65 . 2 92 . 6 94. 4 S plit PP 76. 4 76. 7 76. 5 0. 8 8 60. 2 8 8 . 0 93. 4 S plit VP 72 . 6 71 . 0 71 . 8 0. 8 9 63. 1 8 7. 1 93. 0 S plit S B AR 74. 0 75 . 0 74. 5 0. 70 65 . 4 91 . 1 94. 1 Unary NP 76. 0 76. 4 76. 2 0. 64 65 . 6 92 . 7 94. 3 P P+ NP+ S BAR 77. 7 77. 8 77. 8 0. 94 60. 2 8 6. 8 93. 4
Table 3. 1 1 . S coring effects on the sister-head model ( with perfect tags)
P recision Recall F-score Avg C B 0C B 6 2BC Cov Baseline 68 . 6 66. 9 67. 8 0. 71 65 . 0 8 9. 7 96. 2 P P+ NP+ S BAR 77. 7 77. 8 77. 7 1 . 03 5 8 . 5 8 5 . 1 93. 4
Table 3. 1 2 . S coring effects on the C ollins model ( with perfect tags)
Based upon the performance of these changes on the development set, we also apply the combination of three of the four transformations together. We leave out the VP transformation of Rule ( 3. 6) in this case. This entails performing five experiments: each of the four transformations alone plus one experiment with three of the transformations together. We perform these five experiments on the sister-head parser. For the sake of comparison, we also perform the last experi- ment on the head-head parser.
3 . 5 . 2 Results
The results for the sister-head model summarized in Table 3. 1 1 . The first line shows the results of the baseline model, the sister-head parser without any modifi- cation. ‘ S plit P P’ refers to adding an NP node inside a PP , and this change raises the F-score to 76. 5 from 74. 1 for the baseline. The ‘ S plit VP ’ line shows the result of adding a VP node dominating finite verbs. This transformation causes a dramatic fall in performance, to 71 . 8 . The ‘ S plit S BAR’ op eration provides a moderate improvement of 0. 4 over the baseline. The last individual change, adding ‘ Unary NP’ nodes improved performance to 76. 2 .
The combination of PP splitting, adding unary NP nodes and S BAR splitting improved performance of the sister-head parser to 77. 8 . The effect of the com- bined change on the C ollins model is shown in Table 3. 1 2 . The average number of crossing brackets in the last condition is 1 . 03 with 5 8 . 5 % of sentences having no crossing brackets and 8 5 . 1 % of sentences having no more than two.
3 . 5 . 3 D iscussion
M ost of the unflattening operations helped. Moreover, the difference in perfor- mance of the sister-head and head-head model fell dramatically, from a difference of 8 points to a difference of about 2 points. This justifies the argument that much of the difference in performance between the head-head and sister-head parser in NEG RA are indeed due to assumptions about annotation. In addition, the higher overall scores appear to justify that scoring effects account for some of the differences between scores of the parser in NEGRA and the WS J.
O ne interesting finding is that adding an explicit VP node dominating the finite verb did not help improve overall scores. It has been argued that freer-word order languages have an intrinsically flatter structure, and, in particular, that there is evidence that VP nodes do not exist in G erman . We are agnostic to this theoretical linguistic claim, but it is nonetheless interesting to point out that of all the unflattening operations we attempt, only this one actually hurt perfor- mance.