TY - GEN
T1 - A gold standard dependency corpus for English
AU - Silveira, Natalia
AU - Dozat, Timothy
AU - De Marneffe, Marie Catherine
AU - Bowman, Samuel R.
AU - Connor, Miriam
AU - Bauer, John
AU - Manning, Christopher D.
PY - 2014
Y1 - 2014
N2 - We present a gold standard annotation of syntactic dependencies in the English Web Treebank corpus using the Stanford Dependencies standard. This resource addresses the lack of a gold standard dependency treebank for English, as well as the limited availability of gold standard syntactic annotations for informal genres of English text. We also present experiments on the use of this resource, both for training dependency parsers and for evaluating dependency parsers like the one included as part of the Stanford Parser. We show that training a dependency parser on a mix of newswire and web data improves performance on that type of data without greatly hurting performance on newswire text, and therefore gold standard annotations for non-canonical text can be valuable for parsing in general. Furthermore, the systematic annotation effort has informed both the SD formalism and its implementation in the Stanford Parser's dependency converter. In response to the challenges encountered by annotators in the EWT corpus, we revised and extended the Stanford Dependencies standard, and improved the Stanford Parser's dependency converter.
AB - We present a gold standard annotation of syntactic dependencies in the English Web Treebank corpus using the Stanford Dependencies standard. This resource addresses the lack of a gold standard dependency treebank for English, as well as the limited availability of gold standard syntactic annotations for informal genres of English text. We also present experiments on the use of this resource, both for training dependency parsers and for evaluating dependency parsers like the one included as part of the Stanford Parser. We show that training a dependency parser on a mix of newswire and web data improves performance on that type of data without greatly hurting performance on newswire text, and therefore gold standard annotations for non-canonical text can be valuable for parsing in general. Furthermore, the systematic annotation effort has informed both the SD formalism and its implementation in the Stanford Parser's dependency converter. In response to the challenges encountered by annotators in the EWT corpus, we revised and extended the Stanford Dependencies standard, and improved the Stanford Parser's dependency converter.
KW - Dependency grammar
KW - Stanford dependencies
KW - Web treebank
UR - http://www.scopus.com/inward/record.url?scp=84977556694&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=84977556694&partnerID=8YFLogxK
M3 - Conference contribution
AN - SCOPUS:84977556694
T3 - Proceedings of the 9th International Conference on Language Resources and Evaluation, LREC 2014
SP - 2897
EP - 2904
BT - Proceedings of the 9th International Conference on Language Resources and Evaluation, LREC 2014
A2 - Calzolari, Nicoletta
A2 - Choukri, Khalid
A2 - Goggi, Sara
A2 - Declerck, Thierry
A2 - Mariani, Joseph
A2 - Maegaard, Bente
A2 - Moreno, Asuncion
A2 - Odijk, Jan
A2 - Mazo, Helene
A2 - Piperidis, Stelios
A2 - Loftsson, Hrafn
PB - European Language Resources Association (ELRA)
T2 - 9th International Conference on Language Resources and Evaluation, LREC 2014
Y2 - 26 May 2014 through 31 May 2014
ER -