Irish treebanking and parsing: a preliminary evaluation

Teresa Lynn, Özlem Çetinoğlu, Jennifer Foster, Elaine Uí Dhonnchadha, Mark Dras, Josef van Genabith

Research output: Chapter in Book/Report/Conference proceedingConference proceeding contributionpeer-review

6 Citations (Scopus)


Language resources are essential for linguistic research and the development of NLP applications. Low-density languages, such as Irish, therefore lack significant research in this area. This paper describes the early stages in the development of new language resources for Irish – namely the first Irish dependency treebank and the first Irish statistical dependency parser. We present the methodology behind building our new treebank and the steps we take to leverage upon the few existing resources. We discuss language-specific choices made when defining our dependency labelling scheme, and describe interesting Irish language characteristics such as prepositional attachment, copula and clefting. We manually develop a small treebank of 300 sentences based on an existing POS-tagged corpus and report an inter-annotator agreement of 0.7902. We train MaltParser to achieve preliminary parsing results for Irish and describe a bootstrapping approach for further stages of development.
Original languageEnglish
Title of host publicationProceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)
EditorsNicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis
PublisherEuropean Language Resources Association (ELRA)
Number of pages8
ISBN (Print)9782951740877
Publication statusPublished - 2012
EventInternational Conference on Language Resources and Evaluation (8th : 2012) - Istanbul, Turkey
Duration: 23 May 201225 May 2012


ConferenceInternational Conference on Language Resources and Evaluation (8th : 2012)
CityIstanbul, Turkey


  • dependency
  • treebank
  • Irish


Dive into the research topics of 'Irish treebanking and parsing: a preliminary evaluation'. Together they form a unique fingerprint.

Cite this