diff --git a/_includes/accordion.html b/_includes/accordion.html deleted file mode 100644 index 3a656d6788..0000000000 --- a/_includes/accordion.html +++ /dev/null @@ -1,121 +0,0 @@ - - - diff --git a/_includes/at_glance.html b/_includes/at_glance.html index c0912ea2ae..7802939d1c 100644 --- a/_includes/at_glance.html +++ b/_includes/at_glance.html @@ -1,29 +1,22 @@ - - - - +
+ + + Abaza + 1 + <1K + + + Northwest Caucasian + +
- -
- - -

Abaza treebanks

-
- - - -
+ +
UD_Abaza-ATB is a treebank based on [Spoken corpus of Abaza](http://lingconlab.ru/spoken_abaza/). @@ -45,11 +38,10 @@

Abaza treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -59,35 +51,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Abkhaz + 1 + 13K + + + Northwest Caucasian + +
    -

    Abkhaz treebanks

    -
    - - - -
    + +
    UD_Abkhaz-AbNC is a treebank based on texts from the Abkhaz National Corpus, [AbNC](https://clarino.uib.no/abnc). @@ -109,11 +95,10 @@

    Abkhaz treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -123,35 +108,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Afrikaans + 1 + 49K + + + IE, Germanic + +
    - -
    - - -

    Afrikaans treebanks

    -
    - - - -
    + +
    UD Afrikaans-AfriBooms is a conversion of the AfriBooms Dependency Treebank, originally annotated with a simplified PoS set and dependency relations according to a subset of the Stanford tag set. The corpus consists of public government documents. @@ -173,11 +152,10 @@

    Afrikaans treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -187,35 +165,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Akkadian + 2 + 25K + + + Afro-Asiatic, Semitic + +
    -

    Akkadian treebanks

    -
    - - - -
    + +
    162 royal inscriptions of four early Neo-Assyrian kings. @@ -237,12 +209,11 @@

    Akkadian treebanks

  • Download
  • -

     

    - - - -
    + +
    + PISANDUB 1K @@ -251,8 +222,8 @@

    Akkadian treebanks

    -
    -
    + +
    A small set of sentences from Babylonian royal inscriptions. @@ -264,11 +235,10 @@

    Akkadian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Akkadian treebanks. @@ -280,35 +250,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Akuntsu + 1 + 1K + + + Tupian, Tupari + +
    -

    Akuntsu treebanks

    -
    - - - -
    + +
    UD_Akuntsu-TuDeT is a collection of annotated sentences in <a href="https://glottolog.org/resource/languoid/id/akun1241"> Akuntsú</a>. The sentences stem from the grammatical description by Aragon (2014) and Aragon's field work. Sentence annotation and documentation by Carolina Aragon, Fabrício Ferraz Gerardi, Luana dos Santos. @@ -330,11 +294,10 @@

    Akuntsu treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -344,35 +307,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Albanian + 2 + 4K + + + IE, Albanian + +
    -

    Albanian treebanks

    -
    - - - -
    + +
    The UD-Albanian-STAF (Saarbruecken Treebank of Albanian Fiction) is a treebank of the Albanian language, comprising 202 randomly selected sentences from six fictional books published between 1963 and 2004. @@ -394,12 +351,11 @@

    Albanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + TSA <1K @@ -408,8 +364,8 @@

    Albanian treebanks

    -
    -
    + +
    The UD Albanian Treebank is a small treebank for Standard Albanian, developed within a project framework at Uppsala University. The data was extracted from Wikipedia. @@ -421,11 +377,10 @@

    Albanian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Albanian treebanks. @@ -437,35 +392,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Alemannic + 2 + 21K + + + IE, Germanic + +
    -

    Alemannic treebanks

    -
    - - - -
    + +
    _UD\_Alemannic-UZH_ is a tiny manually annotated treebank of 100 sentences in different Swiss German dialects and a variety of text genres. @@ -487,12 +436,11 @@

    Alemannic treebanks

  • Download
  • -

     

    - - - -
    + +
    + DIVITAL 19K @@ -501,8 +449,8 @@

    Alemannic treebanks

    -
    -
    + +
    UD_Alemannic-DIVITAL is a manually corrected treebank of Alemannic Alsatian consisting of sentences from several genres. @@ -514,11 +462,10 @@

    Alemannic treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Alemannic treebanks. @@ -530,35 +477,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Amharic + 1 + 10K + + + Afro-Asiatic, Semitic + +
    -

    Amharic treebanks

    -
    - - - -
    + +
    UD_Amharic-ATT is a manual developed Treebanks for Amharic. Sentences were collected from grammar books, fictions, biographies, religious texts and news. @@ -580,11 +521,10 @@

    Amharic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -594,35 +534,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Ancient Greek + 3 + 456K + + + IE, Greek + +
    -

    Ancient Greek treebanks

    -
    - - - -
    + +
    UD Ancient Greek PTNK contains portions of the Septuagint according to the Codex Alexandrinus. @@ -644,12 +578,11 @@

    Ancient Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + PROIEL 214K @@ -658,8 +591,8 @@

    Ancient Greek treebanks

    -
    -
    + +
    UD_Ancient_Greek-PROIEL is converted from the Ancient Greek data in the PROIEL treebank, and consists of the New Testament plus selections from Herodotus. @@ -671,12 +604,11 @@

    Ancient Greek treebanks

  • Download
  • -

     

    - - - - -
    + +
    This Universal Dependencies Ancient Greek Treebank consists of an automatic conversion of a selection of passages from the Ancient Greek and Latin @@ -700,11 +632,10 @@

    Ancient Greek treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Ancient Greek treebanks. @@ -716,35 +647,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Ancient Hebrew + 1 + 145K + + + Afro-Asiatic, Semitic + +
    -

    Ancient Hebrew treebanks

    -
    - - - -
    + +
    UD Ancient Hebrew PTNK contains portions of the Biblia Hebraic Stuttgartensia with morphological annotations from [ETCBC](https://github.com/etcbc/bhsa) and syntactic annotations partially based on [MACULA](https://github.com/Clear-Bible/macula-hebrew/). @@ -766,11 +691,10 @@

    Ancient Hebrew treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -780,35 +704,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Apurina + 1 + 1K + + + Arawakan, Purus + +
    -

    Apurina treebanks

    -
    - - - -
    + +
    This is an Apurinã treebank consisting of sentences from a grammatical description of the language by Maília Fernanda. @@ -830,11 +748,10 @@

    Apurina treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -844,35 +761,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Arabic + 3 + 1,042K + + + Afro-Asiatic, Semitic + +
    - -
    - - -

    Arabic treebanks

    -
    - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -896,12 +807,11 @@

    Arabic treebanks

  • Download
  • -

     

    - - - -
    + +
    + PADT 282K @@ -910,8 +820,8 @@

    Arabic treebanks

    - -
    +
    +
    The Arabic-PADT UD treebank is based on the [Prague Arabic Dependency Treebank](http://ufal.mff.cuni.cz/padt/) (PADT), @@ -925,12 +835,11 @@

    Arabic treebanks

  • Download
  • -

     

    - - - -
    + +
    + NYUAD 738K @@ -939,8 +848,8 @@

    Arabic treebanks

    - -
    +
    +
    The NYUAD Arabic UD treebank is based on the Penn Arabic Treebank (PATB), parts 1, 2, and 3, through conversion to CATiB dependency trees. @@ -952,11 +861,10 @@

    Arabic treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Arabic treebanks. @@ -968,35 +876,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Armenian + 2 + 150K + + + IE, Armenian + +
    - -
    - - -

    Armenian treebanks

    -
    - - - -
    + +
    A Universal Dependencies treebank for Eastern Armenian developed for UD originally by the ArmTDP team led by Marat M. Yavrumyan at the Yerevan State University. @@ -1018,12 +920,11 @@

    Armenian treebanks

  • Download
  • -

     

    - - - -
    + +
    + BSUT 46K @@ -1032,8 +933,8 @@

    Armenian treebanks

    - -
    +
    +
    A Universal Dependencies treebank for Eastern Armenian developed for UD originally by the ArmTDP team led by Marat M. Yavrumyan at the V. Brusov State University in Yerevan. @@ -1045,11 +946,10 @@

    Armenian treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Armenian treebanks. @@ -1061,35 +961,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Assamese + 1 + <1K + + + IE, Indic + +
    -

    Assamese treebanks

    -
    - - - -
    + +
    The Assamese-AiW treebank is a manually annotated corpus in Assamese (Assamese script). Assamese is an Indo-Aryan language written in the Assamese script, from Left-to-Right. Word order is Subject-Object-Verb (SOV) with relatively free constituent order. @@ -1111,11 +1005,10 @@

    Assamese treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1125,35 +1018,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Assyrian + 1 + <1K + + + Afro-Asiatic, Semitic + +
    -

    Assyrian treebanks

    -
    - - - -
    + +
    The Uppsala Assyrian Treebank is a small treebank for Modern Standard Assyrian. The corpus is collected and annotated manually. The data was randomly collected from different textbooks and a short translation of The Merchant of Venice. @@ -1175,11 +1062,10 @@

    Assyrian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1189,35 +1075,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Azerbaijani + 1 + <1K + + + Turkic, Southwestern + +
    - -
    - - -

    Azerbaijani treebanks

    -
    - - - -
    + +
    This is a small treebank of grammatical examples for Azerbaijani. The treebank tries to be neutral about the particular variety (North or @@ -1242,11 +1122,10 @@

    Azerbaijani treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1256,35 +1135,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Bambara + 1 + 13K + + + Mande + +
    -

    Bambara treebanks

    -
    - - - -
    + +
    The UD Bambara treebank is a section of the Corpus Référence du Bambara annotated natively with Universal Dependencies. @@ -1306,11 +1179,10 @@

    Bambara treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1320,35 +1192,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Basque + 1 + 121K + + + Basque + +
    -

    Basque treebanks

    -
    - - - -
    + +
    The Basque UD treebank is based on a automatic conversion from part of the Basque Dependency Treebank (BDT), created at the University of of the Basque Country by the IXA NLP research group. The treebank consists of 8.993 sentences (121.443 tokens) and covers mainly literary and journalistic texts. @@ -1370,11 +1236,10 @@

    Basque treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1384,35 +1249,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Bavarian + 1 + 15K + + + IE, Germanic + +
    - -
    - - -

    Bavarian treebanks

    -
    - - - -
    + +
    MaiBaam is manually annotated with part-of-speech tag, syntactic dependencies, and German lemmas. The treebank encompasses diverse text genres (wiki articles and discussions, grammar examples, fiction, and commands for virtual assistants) and dialects from the North, Central and South Bavarian areas as well as the dialectal transition areas in between. @@ -1435,11 +1294,10 @@

    Bavarian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1449,35 +1307,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Beja + 1 + 11K + + + Afro-Asiatic, Cushitic + +
    - -
    - - -

    Beja treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for Beja, North-Cushitic branch of the Afro-Asiatic phylum mainly spoken in Sudan, Egypt and Eritrea. @@ -1499,11 +1351,10 @@

    Beja treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1513,35 +1364,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Belarusian + 1 + 305K + + + IE, Slavic + +
    -

    Belarusian treebanks

    -
    - - - -
    + +
    The Belarusian UD treebank is based on a sample of the news texts included in the Belarusian-Russian parallel subcorpus of the Russian National Corpus, online search available at: http://ruscorpora.ru/search-para-be.html. @@ -1564,11 +1409,10 @@

    Belarusian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1578,35 +1422,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Bengali + 1 + <1K + + + IE, Indic + +
    -

    Bengali treebanks

    -
    - - - -
    + +
    The BRU Bengali treebank has been created at Begum Rokeya University, Rangpur, by the members of Semantics Lab. @@ -1628,11 +1466,10 @@

    Bengali treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1642,35 +1479,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Bhojpuri + 1 + 6K + + + IE, Indic + +
    - -
    - - -

    Bhojpuri treebanks

    -
    - - - -
    + +
    The [Bhojpuri](https://en.wikipedia.org/wiki/Bhojpuri_language) UD Treebank (BHTB) is a part of the [Universal Dependency treebank](http://universaldependencies.org/) project. @@ -1692,11 +1523,10 @@

    Bhojpuri treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1706,35 +1536,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Bokota + 1 + 2K + + + Chibchan, Guaymiic + +
    -

    Bokota treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for Bokota, a member of the Chibchan language family. The language is spoken by about 500 speakers in Panama. The variant of the Bokota treebank is spoken in the Comarca Ngobe-Bugle, at the border of the Veraguas Province, along the Caribbean coast. The other known variant of the language is called Buglere, and is spoken in the province of Chiriqui, in Panama. @@ -1756,11 +1580,10 @@

    Bokota treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1770,35 +1593,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Bororo + 1 + 160K + + + Bororoan + +
    -

    Bororo treebanks

    -
    - - - -
    + +
    UD_Bororo-BDT is a compilation of annotated sentences in [Bororo](https://glottolog.org/resource/languoid/id/boro1282). The corpus encompasses sentences derived from diverse sources: grammar examples, mythological narratives, fieldwork material, and other sources. Sentence annotation and documentation by [Fabrício Ferraz Gerardi](https://languagestructure.github.io). @@ -1820,11 +1637,10 @@

    Bororo treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1834,35 +1650,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Brahui + 1 + <1K + + + Dravidian + +
    - -
    - - -

    Brahui treebanks

    -
    - - - -
    + +
    The Kholum treebank is a manually annotated corpus in Brahui. @@ -1884,11 +1694,10 @@

    Brahui treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1898,35 +1707,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Breton + 1 + 10K + + + IE, Celtic + +
    - -
    - - -

    Breton treebanks

    -
    - - - -
    + +
    UD Breton-KEB is a treebank of Breton that has been manually annotated according to the Universal Dependencies guidelines. The tokenisation guidelines and morphological annotation comes from a finite-state morphological analyser of Breton released as part @@ -1950,11 +1753,10 @@

    Breton treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1964,35 +1766,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Bulgarian + 1 + 156K + + + IE, Slavic + +
    -

    Bulgarian treebanks

    -
    - - - -
    + +
    UD_Bulgarian-BTB is based on the HPSG-based BulTreeBank, created at the Institute of Information and Communication Technologies, @@ -2021,11 +1817,10 @@

    Bulgarian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2035,35 +1830,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Buryat + 1 + 10K + + + Mongolic + +
    -

    Buryat treebanks

    -
    - - - -
    + +
    The UD Buryat treebank was annotated manually natively in UD and contains grammar book sentences, along with news and some fiction. @@ -2086,11 +1875,10 @@

    Buryat treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2100,35 +1888,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Cantonese + 1 + 13K + + + Sino-Tibetan, Chinese + +
    - -
    - - -

    Cantonese treebanks

    -
    - - - -
    + +
    A Cantonese treebank (in Traditional Chinese characters) of film subtitles and of legislative proceedings of Hong Kong, parallel with the Chinese-HK treebank. @@ -2151,11 +1933,10 @@

    Cantonese treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2165,35 +1946,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Cappadocian + 2 + 4K + + + IE, Greek + +
    -

    Cappadocian treebanks

    -
    - - - -
    + +
    The "Asia Minor Greek in Contact" treebank (AMGiC, UD_AMGiC) is compiled from sentences entailing contact-induced morphosyntactic phenomena (CIMSP) that are a result of the contact between Greek and Turkish varieties in Anatolia and in adjacent regions. The sentences are traced in Asia Minor Greek (AMG) dialectal sources. In addition to the UD analysis, the AMGiC treebank provides information concerning the sociolinguistic context within which CIMSP arise. @@ -2215,12 +1990,11 @@

    Cappadocian treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL 4K @@ -2229,8 +2003,8 @@

    Cappadocian treebanks

    -
    -
    + +
    This is a treebank of Pharasiot, a critically endangered Greek dialect originally spoken near Cappadocia. @@ -2244,11 +2018,10 @@

    Cappadocian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Cappadocian treebanks. @@ -2260,35 +2033,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Catalan + 1 + 547K + + + IE, Romance + +
    -

    Catalan treebanks

    -
    - - - -
    + +
    Catalan data from the [AnCora](http://clic.ub.edu/corpus/) corpus. @@ -2310,11 +2077,10 @@

    Catalan treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2324,35 +2090,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Cebuano + 1 + 1K + + + Austronesian, Greater Central Philippine + +
    -

    Cebuano treebanks

    -
    - - - -
    + +
    UD_Cebuano_GJA is a collection of annotated Cebuano sample sentences randomly taken from three different sources: community-contributed samples from the website Tatoeba, a Cebuano grammar book by Bunye & Yap (1971) and Tanangkinsing's reference grammar on Cebuano (2011). This project is currently work in progress. @@ -2374,11 +2134,10 @@

    Cebuano treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2388,35 +2147,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Central Kurdish + 1 + <1K + + + IE, Iranian + +
    -

    Central Kurdish treebanks

    -
    - - - -
    + +
    This treebank contains manually annotated data for Mukri Kurdish (Indo-European) belonging to Kurdish language family, following the Universal Dependencies (UD) guidelines. It aims to offer a syntactically and morphologically consistent dataset that helps with Kurdish language processing and cross-linguistic studies. The current release includes texts in Kurdish Roman Alphabet script and provides dependency annotation at the word, phrase, and sentence levels. @@ -2438,11 +2191,10 @@

    Central Kurdish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2452,35 +2204,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Chinese + 7 + 309K + + + Sino-Tibetan, Chinese + +
    - -
    - - -

    Chinese treebanks

    -
    - - - -
    + +
    Simplified Chinese Universal Dependencies dataset converted from the GSD (traditional) dataset with manual corrections. @@ -2502,12 +2248,11 @@

    Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + GSD 123K @@ -2516,8 +2261,8 @@

    Chinese treebanks

    - -
    +
    +
    Traditional Chinese Universal Dependencies Treebank annotated and converted by Google. @@ -2530,12 +2275,11 @@

    Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + Beginner 19K @@ -2544,8 +2288,8 @@

    Chinese treebanks

    - -
    +
    +
    A treebank of Chinese sentences adapted for learner of level A1 to C1 (HSK1 to 5) collected on the [Chinese Grammar Wiki](https://resources.allsetlearning.com/chinese/grammar/\) (CC BY-NC-SA 3.0 License) website. The treebank was manually annotated by researchers of Paris Nanterre University (Modyco) in the mSUD annotation schema (morpheme level Surface Universal Dependencies). @@ -2557,12 +2301,11 @@

    Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 21K @@ -2571,8 +2314,8 @@

    Chinese treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -2586,12 +2329,11 @@

    Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + HK 9K @@ -2600,8 +2342,8 @@

    Chinese treebanks

    - -
    +
    +
    A Traditional Chinese treebank of film subtitles and of legislative proceedings of Hong Kong, parallel with the Cantonese-HK treebank. @@ -2614,12 +2356,11 @@

    Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + CFL 7K @@ -2628,8 +2369,8 @@

    Chinese treebanks

    - -
    +
    +
    The Chinese-CFL UD treebank is manually annotated by Keying Li with minor manual revisions by Herman Leung and John Lee at City University of Hong Kong, based on essays written by learners of Mandarin Chinese as a foreign language. The data is in Simplified Chinese. @@ -2641,12 +2382,11 @@

    Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + PatentChar 4K @@ -2655,8 +2395,8 @@

    Chinese treebanks

    - -
    +
    +
    A treebank of Chinese patent application texts collected from the Chinese patent office's website CNIPA. @@ -2670,11 +2410,10 @@

    Chinese treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Chinese treebanks. @@ -2686,35 +2425,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Chintang + 1 + 14K + + + Sino-Tibetan, Himalayish + +
    -

    Chintang treebanks

    -
    - - - -
    + +
    UD\_Chintang-CTNTB is a Universal Dependencies (UD) treebank for the Chintang language. The annotation converted from glosses from "A Grammar of Chintang A Tibeto-Burman Language of Nepal" by Robert Schikowski. @@ -2736,11 +2469,10 @@

    Chintang treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2750,35 +2482,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Chukchi + 1 + 6K + + + Chukotko-Kamchatkan + +
    -

    Chukchi treebanks

    -
    - - - -
    + +
    This data is a manual annotation of the corpus from multimedia annotated corpus of the [Chuklang](http://chuklang.ru/) project, a dialectal corpus of the Amguema @@ -2802,11 +2528,10 @@

    Chukchi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2816,35 +2541,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Classical Armenian + 1 + 99K + + + IE, Armenian + +
    - -
    - - -

    Classical Armenian treebanks

    -
    - - - -
    + +
    The present release includes the Classical Armenian translation of the Gospels and the first book of the "History of the Armenians" by Movses Khorenatsi. The annotation of the Gospels results from a rule-based conversion from the PROIEL annotation, manually corrected and extended with additional information. The annotation of the "History of the Armenians" has been performed by a UDPipe2 annotator and manually corrected. @@ -2866,11 +2585,10 @@

    Classical Armenian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2880,35 +2598,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Classical Chinese + 2 + 433K + + + Sino-Tibetan, Chinese + +
    - -
    - - -

    Classical Chinese treebanks

    -
    - - - -
    + +
    Classical Chinese Universal Dependencies Treebank annotated and converted by Institute for Research in Humanities, Kyoto University. @@ -2930,12 +2642,11 @@

    Classical Chinese treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL <1K @@ -2944,8 +2655,8 @@

    Classical Chinese treebanks

    - -
    +
    +
    A dependency Treebank of "逍遥游(Enjoyment in Untroubled Ease)" written by Zhuangzi. @@ -2957,11 +2668,10 @@

    Classical Chinese treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Classical Chinese treebanks. @@ -2973,35 +2683,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Coptic + 2 + 91K + + + Afro-Asiatic, Egyptian + +
    - -
    - - -

    Coptic treebanks

    -
    - - - -
    + +
    UD Coptic contains manually annotated Sahidic Coptic texts, including Biblical texts, sermons, letters, and hagiography. @@ -3023,12 +2727,11 @@

    Coptic treebanks

  • Download
  • -

     

    - - - -
    + +
    + Bohairic 32K @@ -3037,8 +2740,8 @@

    Coptic treebanks

    - -
    +
    +
    UD_Coptic-Bohairic contains manually annotated Bohairic Coptic texts, including Biblical narrative and poetic texts, epistles, and hagiography. @@ -3050,11 +2753,10 @@

    Coptic treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Coptic treebanks. @@ -3066,35 +2768,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Croatian + 1 + 199K + + + IE, Slavic + +
    - -
    - - -

    Croatian treebanks

    -
    - - - -
    + +
    The Croatian UD treebank is based on the extension of the SETimes-HR corpus, the [hr500k](http://hdl.handle.net/11356/1183) corpus. @@ -3116,11 +2812,10 @@

    Croatian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3130,35 +2825,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Czech + 6 + 4,162K + + + IE, Slavic + +
    -

    Czech treebanks

    -
    - - - -
    + +
    The Czech-PDTC UD treebank is based on the Prague Dependency Treebank – Consolidated (PDT-C) 2.0, created at the Charles University in Prague. @@ -3181,12 +2870,11 @@

    Czech treebanks

  • Download
  • -

     

    - - - -
    + +
    + CAC 494K @@ -3195,8 +2883,8 @@

    Czech treebanks

    -
    -
    + +
    The UD_Czech-CAC treebank is based on the Czech Academic Corpus 2.0 (CAC; Český akademický korpus; ČAK), created at Charles University in Prague. @@ -3209,12 +2897,11 @@

    Czech treebanks

  • Download
  • -

     

    - - - - -
    + +
    FicTree is a treebank of Czech fiction, automatically converted into the UD format. The treebank was built at Charles University in Prague. @@ -3237,12 +2924,11 @@

    Czech treebanks

  • Download
  • -

     

    - - - - -
    + +
    The UD_Czech-CLTT treebank is based on the Czech Legal Text Treebank 2.0, created at the Charles University in Prague. @@ -3265,12 +2951,11 @@

    Czech treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -3294,12 +2979,11 @@

    Czech treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD_Czech-Poetry contains random samples of Czech 19th-century poetry from the Corpus of Czech Verse parsed with UDPipe2 (trained on UD Czech-PDT 2.11) and manually corrected. @@ -3321,11 +3005,10 @@

    Czech treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Czech treebanks. @@ -3337,35 +3020,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Danish + 1 + 100K + + + IE, Germanic + +
    - -
    - - -

    Danish treebanks

    -
    - - - -
    + +
    The Danish UD treebank is a conversion of the Danish Dependency Treebank. @@ -3387,11 +3064,10 @@

    Danish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3401,35 +3077,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - - - -
    - +
    + + + Dutch + 2 + 505K + + + IE, Germanic + +
    -

    Dutch treebanks

    -
    - - - -
    + +
    This corpus contains sentences from the Wikipedia section of the Lassy Small Treebank. Universal Dependency annotation was generated automatically from the original annotation in Lassy. @@ -3452,12 +3122,11 @@

    Dutch treebanks

  • Download
  • -

     

    - - - -
    + +
    + Alpino 208K @@ -3466,8 +3135,8 @@

    Dutch treebanks

    -
    -
    + +
    This corpus consists of samples from various treebanks annotated at the University of Groningen using the Alpino annotation tools and guidelines. @@ -3479,11 +3148,10 @@

    Dutch treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Dutch treebanks. @@ -3495,35 +3163,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Egyptian + 1 + 34K + + + Afro-Asiatic, Egyptian + +
    -

    Egyptian treebanks

    -
    - - - -
    + +
    Egyptian-PC is the first dependency treebank created for the morphosyntactic annotation of pre-Coptic Egyptian. It is developed at the University of Jaén. Its current state (UD v2.18) consists of 3,089 sentences and 34,234 tokens manually annotated from the Pyramid Texts. @@ -3545,11 +3207,10 @@

    Egyptian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3559,35 +3220,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + English + 13 + 1,126K + + + IE, Germanic + +
    -

    English treebanks

    -
    - - - -
    + +
    Universal Dependencies syntax annotations from the GUM corpus (https://gucorpling.org/gum/) @@ -3609,12 +3264,11 @@

    English treebanks

  • Download
  • -

     

    - - - -
    + +
    + EWT 254K @@ -3623,8 +3277,8 @@

    English treebanks

    -
    -
    + +
    A Gold Standard Universal Dependencies Corpus for English, built over the source material of the English Web Treebank @@ -3638,12 +3292,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD English_LinES is the English half of the LinES Parallel Treebank with the original dependency annotation first automatically converted @@ -3668,12 +3321,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD_English-ParTUT is a conversion of a multilingual parallel treebank developed at the University of Turin, and consisting of a variety of text genres, including talks, legal texts and Wikipedia articles, among others. @@ -3696,12 +3348,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD Atis Treebank is a manually annotated treebank consisting of the sentences in the Atis (Airline Travel Informations) dataset which includes the human speech transcriptions of people asking for flight information on the automated inquiry systems. @@ -3723,12 +3374,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    Repository for the Genre Tests for Linguistic Evaluation (GENTLE) Corpus @@ -3750,12 +3400,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    This repository contains Universal Dependencies (UD) trees for utterances from child–adult spoken interactions in English, drawn from [CHILDES](https://childes.talkbank.org/) transcripts. @@ -3777,12 +3426,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    This treebank contains manually corrected Universal Dependency annotations for 500 sentences from the English translation of *The Little Prince*. @@ -3804,12 +3452,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is the English portion of the Parallel Universal Dependencies (PUD) treebanks created for the CoNLL 2017 shared task on Multilingual Parsing @@ -3834,12 +3481,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD_English-CTeTex is a technical text corpus annotated in Universal Dependency syntax containing 196 software requirements. @@ -3861,12 +3507,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD English-Pronouns is dataset created to make pronoun identification more accurate and with a more balanced distribution across genders. The dataset is initially targeting the Independent Genitive pronouns, "hers", (independent) "his", (singular) "theirs", "mine", and (singular) "yours". @@ -3888,12 +3533,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    Universal Dependencies syntax annotations from the Reddit portion of the GUM corpus (https://gucorpling.org/gum/) @@ -3915,12 +3559,11 @@

    English treebanks

  • Download
  • -

     

    - - - - -
    + +
    This repository includes the Dependency Treebank of Spoken L2 English (SL2E), which consists of Universal Dependency annotations for a random sample of sentences from the <a href="https://alaginrc.nict.go.jp/nict_jle/index_E.html" target="_blank">NICT JLE</a>, a corpus of spoken second language English. <a href="https://github.com/LCR-ADS-Lab/SL2E-Dependency-Treebank" target="_blank">The homepage of the project is here.</a> @@ -3942,11 +3585,10 @@

    English treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of English treebanks. @@ -3958,35 +3600,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - - - -
    - +
    + + + Erzya + 1 + 20K + + + Uralic, Mordvin + +
    -

    Erzya treebanks

    -
    - - - -
    + +
    UD Erzya is the original annotation (CoNLL-U) for texts in the Erzya language, it originally consists of a sample from a number of fiction authors writing originals in Erzya. @@ -4009,11 +3645,10 @@

    Erzya treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4023,35 +3658,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Esperanto + 2 + 3K + + + Constructed + +
    - -
    - - -

    Esperanto treebanks

    -
    - - - -
    + +
    UD Esperanto-Prago is the Universal Dependencies syntax annotation on Manifesto de Prago (Prague Manifesto) and Deklaratio pri Homaranismo. @@ -4073,12 +3702,11 @@

    Esperanto treebanks

  • Download
  • -

     

    - - - -
    + +
    + Cairo <1K @@ -4087,8 +3715,8 @@

    Esperanto treebanks

    -
    -
    + +
    This is an example treebank made to ilustrate UD annotation choices made for Esperanto based on the Cairo sample sentences. @@ -4100,11 +3728,10 @@

    Esperanto treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Esperanto treebanks. @@ -4116,35 +3743,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - - - -
    - +
    + + + Estonian + 2 + 528K + + + Uralic, Finnic + +
    -

    Estonian treebanks

    -
    - - - -
    + +
    UD Estonian is a converted version of the Estonian Dependency Treebank (EDT), originally annotated in the Constraint Grammar (CG) annotation scheme, and consisting of genres of fiction, newspaper texts and scientific texts. The treebank contains 30,972 trees, 437,769 tokens. @@ -4166,12 +3787,11 @@

    Estonian treebanks

  • Download
  • -

     

    - - - -
    + +
    + EWT 90K @@ -4180,8 +3800,8 @@

    Estonian treebanks

    -
    -
    + +
    UD EWT treebank consists of different genres of new media. The treebank contains 7,190 trees, 90,585 tokens. @@ -4193,11 +3813,10 @@

    Estonian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Estonian treebanks. @@ -4209,35 +3828,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Faroese + 2 + 50K + + + IE, Germanic + +
    - -
    - - -

    Faroese treebanks

    -
    - - - -
    + +
    This is a treebank of Faroese based on the Faroese Wikipedia. @@ -4259,12 +3872,11 @@

    Faroese treebanks

  • Download
  • -

     

    - - - -
    + +
    + FarPaHC 40K @@ -4273,8 +3885,8 @@

    Faroese treebanks

    -
    -
    + +
    UD_Faroese-FarPaHC is a conversion of the [Faroese Parsed Historical Corpus (FarPaHC)](https://github.com/einarfs/farpahc) to the Universal Dependencies scheme. @@ -4288,11 +3900,10 @@

    Faroese treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Faroese treebanks. @@ -4304,35 +3915,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - - - -
    - +
    + + + Finnish + 4 + 397K + + + Uralic, Finnic + +
    -

    Finnish treebanks

    -
    - - - -
    + +
    UD_Finnish-TDT is based on the Turku Dependency Treebank (TDT), a broad-coverage dependency treebank of general Finnish covering numerous genres. The conversion to UD was followed by extensive manual checks and corrections, and the treebank closely adheres to the UD guidelines. @@ -4354,12 +3959,11 @@

    Finnish treebanks

  • Download
  • -

     

    - - - -
    + +
    + FTB 159K @@ -4368,8 +3972,8 @@

    Finnish treebanks

    -
    -
    + +
    FinnTreeBank 1 consists of manually annotated grammatical examples from VISK. The UD version of FinnTreeBank 1 was converted from a @@ -4383,12 +3987,11 @@

    Finnish treebanks

  • Download
  • -

     

    - - - - -
    + +
    Finnish-OOD is an external out-of-domain test set for Finnish-TDT annotated natively into UD scheme. @@ -4410,12 +4013,11 @@

    Finnish treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -4439,11 +4041,10 @@

    Finnish treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Finnish treebanks. @@ -4455,35 +4056,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + French + 9 + 708K + + + IE, Romance + +
    - -
    - - -

    French treebanks

    -
    - - - -
    + +
    The **UD_French-GSD** was converted in 2015 from the content head version of the universal dependency treebank v2.0 (https://github.com/ryanmcd/uni-dep-tb). @@ -4507,12 +4102,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + ALTS 68K @@ -4521,8 +4115,8 @@

    French treebanks

    - -
    +
    +
    ALTS (AUTOMATED Sixteenth-century corpus) is a treebank of sixteenth-century legal French from Normandy and the Channel Islands. @@ -4534,12 +4128,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + Sequoia 70K @@ -4548,8 +4141,8 @@

    French treebanks

    - -
    +
    +
    **UD_French-Sequoia** is an automatic conversion of the [SUD_French-Sequoia](https://github.com/surfacesyntacticud/SUD_French-Sequoia) treebank, which comes from the former corpus [French Sequoia corpus](http://deep-sequoia.inria.fr). @@ -4561,12 +4154,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + ParisStories 42K @@ -4575,8 +4167,8 @@

    French treebanks

    - -
    +
    +
    Paris Stories is a corpus of oral French collected and transcribed by Linguistics students from Sorbonne Nouvelle and corrected by students from the Plurital Master's Degree of Computational Linguistics ( Inalco, Paris Nanterre, Sorbonne Nouvelle) between 2017 and 2021. It contains monologues and dialogues from speakers living in the Parisian region. @@ -4589,12 +4181,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + Rhapsodie 44K @@ -4603,8 +4194,8 @@

    French treebanks

    - -
    +
    +
    A Universal Dependencies corpus for spoken French. @@ -4616,12 +4207,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + ParTUT 28K @@ -4630,8 +4220,8 @@

    French treebanks

    - -
    +
    +
    UD_French-ParTUT is a conversion of a multilingual parallel treebank developed at the University of Turin, and consisting of a variety of text genres, including talks, legal texts and Wikipedia articles, among others. @@ -4644,12 +4234,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 24K @@ -4658,8 +4247,8 @@

    French treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -4673,12 +4262,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + FQB 23K @@ -4687,8 +4275,8 @@

    French treebanks

    - -
    +
    +
    The corpus **UD_French-FQB** is an automatic conversion of the [French QuestionBank v1](http://alpage.inria.fr/Treebanks/FQB/), a corpus entirely made of questions. @@ -4700,12 +4288,11 @@

    French treebanks

  • Download
  • -

     

    - - - -
    + +
    + PoitevinDIVITAL 5K @@ -4714,8 +4301,8 @@

    French treebanks

    - -
    +
    +
    UD_French-PoitevinDIVITAL is a manually corrected treebank of Poitevin-Saintongeais consisting of sentences from several genres. @@ -4727,11 +4314,10 @@

    French treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of French treebanks. @@ -4743,35 +4329,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Frisian Dutch + 1 + 3K + + + Code switching + +
    -

    Frisian Dutch treebanks

    -
    - - - -
    + +
    UD_Frisian_Dutch-Fame is a selection of 400 sentences from the FAME! speech corpus by Yilmaz et al. (2016a, 2016b). The treebank is manually annotated using the UD scheme. @@ -4793,11 +4373,10 @@

    Frisian Dutch treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4807,35 +4386,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Galician + 3 + 188K + + + IE, Romance + +
    - -
    - - -

    Galician treebanks

    -
    - - - -
    + +
    The Galician-TreeGal is a treebank for Galician developed at LyS Group (Universidade da Coruña) and at CiTIUS (Universidade de Santiago de Compostela). @@ -4857,12 +4430,11 @@

    Galician treebanks

  • Download
  • -

     

    - - - -
    + +
    + CTG 139K @@ -4871,8 +4443,8 @@

    Galician treebanks

    - -
    +
    +
    The Galician UD treebank is based on the automatic parsing of the Galician Technical Corpus (http://sli.uvigo.gal/CTG) created at the University of Vigo by the the TALG NLP research group. @@ -4884,12 +4456,11 @@

    Galician treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 23K @@ -4898,8 +4469,8 @@

    Galician treebanks

    - -
    +
    +
    The Galician PUD is a treebank for Galician developed at CiTIUS (Universidade de Santiago de Compostela). It follows the annotation guidelines of [Galician-TreeGal](https://github.com/UniversalDependencies/UD_Galician-TreeGal). @@ -4911,11 +4482,10 @@

    Galician treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Galician treebanks. @@ -4927,35 +4497,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Georgian + 2 + 83K + + + Kartvelian + +
    -

    Georgian treebanks

    -
    - - - -
    + +
    UD_Georgian-GNC is a treebank based on texts from the Georgian National Corpus, [GNC](https://clarino.uib.no/gnc). @@ -4977,12 +4541,11 @@

    Georgian treebanks

  • Download
  • -

     

    - - - -
    + +
    + GLC 60K @@ -4991,8 +4554,8 @@

    Georgian treebanks

    -
    -
    + +
    The Georgian UD Treebank (UD_Georgian-GLC) is the first syntactically annotated corpus of Georgian, based on a collection of annotated sentences selected from the Georgian Language Corpus (GLC) available at http://corpora.iliauni.edu.ge/ and sentences selected from Wiki in accordance with the 132 scientific fields. @@ -5004,11 +4567,10 @@

    Georgian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Georgian treebanks. @@ -5020,35 +4582,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + German + 4 + 3,810K + + + IE, Germanic + +
    - -
    - - -

    German treebanks

    -
    - - - -
    + +
    UD German-HDT is a conversion of the Hamburg Dependency Treebank, created at the University of Hamburg through manual annotation in conjunction with a standard for morphologically and syntactically annotating sentences as well as a constraint-based parser. @@ -5070,12 +4626,11 @@

    German treebanks

  • Download
  • -

     

    - - - -
    + +
    + GSD 292K @@ -5084,8 +4639,8 @@

    German treebanks

    - -
    +
    +
    The German UD is converted from the content head version of the [universal dependency treebank v2.0 (legacy)](https://github.com/ryanmcd/uni-dep-tb). @@ -5098,12 +4653,11 @@

    German treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 21K @@ -5112,8 +4666,8 @@

    German treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -5127,12 +4681,11 @@

    German treebanks

  • Download
  • -

     

    - - - -
    + +
    + LIT 40K @@ -5141,8 +4694,8 @@

    German treebanks

    - -
    +
    +
    This treebank aims at gathering texts of the German literary history. Currently, it hosts Fragments of the early Romanticism, i.e. aphorism-like texts mainly dealing with philosophical issues concerning art, beauty and related topics. @@ -5154,11 +4707,10 @@

    German treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of German treebanks. @@ -5170,35 +4722,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Gheg + 1 + 15K + + + IE, Albanian + +
    -

    Gheg treebanks

    -
    - - - -
    + +
    UD Gheg Pear Stories (GPS) contains renarrations of Wallace Chafe's Pear Stories video (pearstories.org) by heritage speakers of Gheg Albanian living in Switzerland and speakers from Prishtina. @@ -5220,11 +4766,10 @@

    Gheg treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5234,35 +4779,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Gorontalo + 1 + <1K + + + Austronesian, Greater Central Philippine + +
    - -
    - - -

    Gorontalo treebanks

    -
    - - - -
    + +
    Bungo lo Lombi is a Universal Dependencies parsed corpus of modern spoken Gorontalo as spoken in Gorontalo City, Gorontalo Province, Indonesia. It comprises fieldwork samples obtained by Colleen Alena O'Brien. @@ -5284,11 +4823,10 @@

    Gorontalo treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5298,35 +4836,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Gothic + 1 + 55K + + + IE, Germanic + +
    -

    Gothic treebanks

    -
    - - - -
    + +
    The UD Gothic treebank is based on the Gothic data from the PROIEL treebank, and consists of Wulfila's Bible translation. @@ -5348,11 +4880,10 @@

    Gothic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5362,35 +4893,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Greek + 6 + 110K + + + IE, Greek + +
    - -
    - - -

    Greek treebanks

    -
    - - - -
    + +
    The Greek UD treebank (UD_Greek-GDT) is derived from the Greek Dependency Treebank (https://gdt.ilsp.gr), a resource developed and maintained by @@ -5415,12 +4940,11 @@

    Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + GUD 25K @@ -5429,8 +4953,8 @@

    Greek treebanks

    - -
    +
    +
    GUD is a resource for EL manually annotated for morphology and syntax. It is an ongoing project led by Stella Markantonatou and Vivian Stamou (hereinafter: the GUD team), both researchers at the [Institute for Language and Speech Processing](http://www.ilsp.gr/) (ILSP/Athena Research Centre). @@ -5442,12 +4966,11 @@

    Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + Lesbian 6K @@ -5456,8 +4979,8 @@

    Greek treebanks

    - -
    +
    +
    A Universal Dependencies (UD) treebank for the dialect of Lesbos, a low-resource living Northern variety of Modern Greek. The treebank currently contains 625 sentences with manual annotations following the Universal Dependencies framework, representing the first UD treebank for a Northern Modern Greek dialect. @@ -5469,12 +4992,11 @@

    Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + GLCII 9K @@ -5483,8 +5005,8 @@

    Greek treebanks

    - -
    +
    +
    A treebank based on version 2 of the Greek Learner Corpus (GLCII), consisting of written data produced by learners of Modern Greek. @@ -5496,12 +5018,11 @@

    Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + Messinian <1K @@ -5510,8 +5031,8 @@

    Greek treebanks

    - -
    +
    +
    Messenian is in the Southern group of dialects of Modern Greek to which also belongs the main variety (the Standard). @@ -5523,12 +5044,11 @@

    Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + Cretan 4K @@ -5537,8 +5057,8 @@

    Greek treebanks

    - -
    +
    +
    The text of the treebank was transcribed with Wisper (trained on Cretan) from 9 tapes containing folklore narratives by one speaker, Ioannis Anagnostakis, who is responsible for their composition. The narratives are radio broadcasts in digital format, with permission from the Audiovisual Department of the Vikelaia Municipal Library of Heraklion, Crete (1998-2001). The data were split into training (70%), dev (10%) and test (20%) sets. @@ -5551,11 +5071,10 @@

    Greek treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Greek treebanks. @@ -5567,35 +5086,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Guajajara + 1 + 9K + + + Tupian, Maweti-Guarani + +
    -

    Guajajara treebanks

    -
    - - - -
    + +
    UD_Guajajara-TuDeT is a collection of annotated sentences in <a href="https://glottolog.org/resource/languoid/id/guaj1255">Guajajara</a>. Sentences stem from multiple sources such as descriptions of the language, short stories, dictionaries and translations from the New Testament. Sentence annotation and documentation by Lorena Martín Rodríguez and Fabrício Ferraz Gerardi. @@ -5617,11 +5130,10 @@

    Guajajara treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5631,35 +5143,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Guarani + 1 + <1K + + + Tupian, Maweti-Guarani + +
    - -
    - - -

    Guarani treebanks

    -
    - - - -
    + +
    UD_Guarani-OldTuDeT is a collection of annotated texts in <a href="https://glottolog.org/resource/languoid/id/oldp1258">Old Guaraní</a>. All known sources in this language are being annotated: cathesisms, grammars (seventeenth and eighteenth century), sentences from dictionaries, and other texts. Sentence annotation and documentation by Fabrício Ferraz Gerardi and Lorena Martín Rodríguez. @@ -5681,11 +5187,10 @@

    Guarani treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5695,35 +5200,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Gujarati + 1 + 1K + + + IE, Indic + +
    -

    Gujarati treebanks

    -
    - - - -
    + +
    GujTB is an in-progress treebank of Gujarati (an Indo-Aryan language) in Gujarati script. @@ -5745,11 +5244,10 @@

    Gujarati treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5759,35 +5257,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Gwichin + 1 + 1K + + + Na-Dene + +
    - -
    - - -

    Gwichin treebanks

    -
    - - - -
    + +
    UD_Gwichin-TueCL is a small treebank of Alaskan Gwich'in, an endangered Athabascan language, based on material located in the Alaska Native Language Archive. @@ -5809,11 +5301,10 @@

    Gwichin treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5823,35 +5314,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Haitian Creole + 2 + 75K + + + Creole + +
    -

    Haitian Creole treebanks

    -
    - - - -
    + +
    This is a treebank of Haitian creole. It contains 144 sentences selected from 3 major genres: bible, literary texts, newspapers. @@ -5875,12 +5360,11 @@

    Haitian Creole treebanks

  • Download
  • -

     

    - - - -
    + +
    + Adolphe 71K @@ -5889,8 +5373,8 @@

    Haitian Creole treebanks

    -
    -
    + +
    This is a treebank for Haitian creole. It contains 3314 sentences and 300,000+ words selected from 1 bible-related source and was annotated programmatically. @@ -5904,11 +5388,10 @@

    Haitian Creole treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Haitian Creole treebanks. @@ -5920,35 +5403,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Hausa + 4 + 53K + + + Afro-Asiatic, West Chadic + +
    - -
    - - -

    Hausa treebanks

    -
    - - - -
    + +
    This treebank contains data of the Autogramm project, for the (Kano) Eastern dialect of Hausa, Nigeria. @@ -5970,12 +5447,11 @@

    Hausa treebanks

  • Download
  • -

     

    - - - -
    + +
    + WesternAutogramm 13K @@ -5984,8 +5460,8 @@

    Hausa treebanks

    - -
    +
    +
    This treebank contains data of Southern Autogramm, for the (Tibiri) Gobir dialect of Niger Republic (Western Hausa). @@ -5997,12 +5473,11 @@

    Hausa treebanks

  • Download
  • -

     

    - - - -
    + +
    + NorthernAutogramm 15K @@ -6011,8 +5486,8 @@

    Hausa treebanks

    - -
    +
    +
    This treebank contains data of Northern Autogramm, for the Ader dialect of Niger Republic (Northern Hausa). @@ -6024,12 +5499,11 @@

    Hausa treebanks

  • Download
  • -

     

    - - - -
    + +
    + SouthernAutogramm 14K @@ -6038,8 +5512,8 @@

    Hausa treebanks

    - -
    +
    +
    This treebank contains data of Southern Autogramm, for the Zaria dialect of Nigeria (Southern Hausa). @@ -6051,11 +5525,10 @@

    Hausa treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Hausa treebanks. @@ -6067,35 +5540,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Hebrew + 4 + 376K + + + Afro-Asiatic, Semitic + +
    -

    Hebrew treebanks

    -
    - - - -
    + +
    Publicly available subset of the IAHLT UD Hebrew Treebank's Wikipedia section (https://www.iahlt.org/) @@ -6117,12 +5584,11 @@

    Hebrew treebanks

  • Download
  • -

     

    - - - -
    + +
    + IAHLTknesset 67K @@ -6131,8 +5597,8 @@

    Hebrew treebanks

    -
    -
    + +
    Publicly available IAHLT UD Hebrew Treebank's Knesset section (https://www.iahlt.org/) @@ -6144,12 +5610,11 @@

    Hebrew treebanks

  • Download
  • -

     

    - - - - -
    + +
    A Universal Dependencies Corpus for Hebrew. @@ -6171,12 +5636,11 @@

    Hebrew treebanks

  • Download
  • -

     

    - - - - -
    + +
    A Universal Dependencies treebank of post-Rabbinic historical Hebrew, comprising ~300 (~8000 tokens) sentences annotated for morphology and syntax from diverse pre-modern sources. @@ -6198,11 +5662,10 @@

    Hebrew treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Hebrew treebanks. @@ -6214,35 +5677,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Highland P. Nahuatl + 1 + 10K + + + Uto-Aztecan + +
    - -
    - - -

    Highland Puebla Nahuatl treebanks

    -
    - - - -
    + +
    UD_Highland_Puebla_Nahuatl-ITML is a collection of texts in the Highland Puebla variety of Nahuatl (ISO-639: `azz`) spoken in 24 municipalities in the state of Mexico in Puebla. The treebank contains spoken monologue and dialogue, scientific texts translated from Spanish and some miscellaneous @@ -6266,11 +5723,10 @@

    Highland Puebla Nahuatl treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -6280,35 +5736,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Hindi + 2 + 375K + + + IE, Indic + +
    -

    Hindi treebanks

    -
    - - - -
    + +
    The Hindi UD treebank is based on the Hindi Dependency Treebank (HDTB), created at IIIT Hyderabad, India. @@ -6331,12 +5781,11 @@

    Hindi treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 23K @@ -6345,8 +5794,8 @@

    Hindi treebanks

    -
    -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -6360,11 +5809,10 @@

    Hindi treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Hindi treebanks. @@ -6376,35 +5824,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Hittite + 1 + 1K + + + IE, Anatolian + +
    - -
    - - -

    Hittite treebanks

    -
    - - - -
    + +
    UD_Hittite-HitTB is a small Universal Dependencies treebank for Hittite, containing original sentences from Hoffner and Melchert's tutorial to A Grammar of the Hittite Language. @@ -6426,11 +5868,10 @@

    Hittite treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -6440,35 +5881,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Hungarian + 1 + 42K + + + Uralic, Ugric + +
    -

    Hungarian treebanks

    -
    - - - -
    + +
    The Hungarian UD treebank is derived from the Szeged Dependency Treebank (Vincze et al. 2010). @@ -6490,11 +5925,10 @@

    Hungarian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -6504,35 +5938,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Icelandic + 4 + 1,183K + + + IE, Germanic + +
    - -
    - - -

    Icelandic treebanks

    -
    - - - -
    + +
    UD_Icelandic-IcePaHC is a conversion of the [Icelandic Parsed Historical Corpus (IcePaHC)](https://linguist.is/icelandic_treebank/Icelandic_Parsed_Historical_Corpus_(IcePaHC)) to the Universal Dependencies scheme. @@ -6556,12 +5984,11 @@

    Icelandic treebanks

  • Download
  • -

     

    - - - -
    + +
    + Modern 80K @@ -6570,8 +5997,8 @@

    Icelandic treebanks

    - -
    +
    +
    UD_Icelandic-Modern is a conversion of the [modern additions](https://github.com/antonkarl/icecorpus/tree/master/additions2019) to the Icelandic Parsed Historical Corpus (IcePaHC) to the Universal Dependencies scheme. @@ -6583,12 +6010,11 @@

    Icelandic treebanks

  • Download
  • -

     

    - - - -
    + +
    + GC 99K @@ -6597,8 +6023,8 @@

    Icelandic treebanks

    - -
    +
    +
    UD_Icelandic-GC is a conversion of the gold part of [GreynirCorpus](https://github.com/mideind/GreynirCorpus), which has been manually corrected and verified. The corpus is parsed into full constituency trees, and converted using [UDConverter-GreynirCorpus](https://github.com/thorunna/UDConverter-GreynirCorpus). @@ -6610,12 +6036,11 @@

    Icelandic treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 18K @@ -6624,8 +6049,8 @@

    Icelandic treebanks

    - -
    +
    +
    Icelandic-PUD is the Icelandic part of the Parallel Universal Dependencies (PUD) treebanks. @@ -6637,11 +6062,10 @@

    Icelandic treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Icelandic treebanks. @@ -6653,35 +6077,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Ika + 1 + 5K + + + Chibchan, Arhuacic + +
    -

    Ika treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for Ika, a member of the Chibchan language family. The language is spoken by about 25,000 speakers in Colombia. @@ -6703,11 +6121,10 @@

    Ika treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -6717,35 +6134,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Indonesian + 3 + 169K + + + Austronesian, Malayo-Sumbawan + +
    - -
    - - -

    Indonesian treebanks

    -
    - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -6769,12 +6180,11 @@

    Indonesian treebanks

  • Download
  • -

     

    - - - -
    + +
    + GSD 122K @@ -6783,8 +6193,8 @@

    Indonesian treebanks

    - -
    +
    +
    The Indonesian-GSD treebank was originally converted from the content head version of the [universal dependency treebank v2.0 (legacy)](https://github.com/ryanmcd/uni-dep-tb) in 2015. In order to comply with the latest Indonesian annotation guidelines, the treebank has undergone a major revision between UD releases v2.8 and v2.9 (2021). @@ -6796,12 +6206,11 @@

    Indonesian treebanks

  • Download
  • -

     

    - - - -
    + +
    + CSUI 28K @@ -6810,8 +6219,8 @@

    Indonesian treebanks

    - -
    +
    +
    UD Indonesian-CSUI is a conversion from an Indonesian constituency treebank in the Penn Treebank format named [**Kethu**](https://github.com/ialfina/kethu) that was also a conversion from a constituency treebank built by [**Dinakaramani et al. (2015)**](https://github.com/famrashel/idn-treebank). We named this treebank **Indonesian-CSUI**, since all the three versions of the treebanks were built at Faculty of Computer Science, Universitas Indonesia. @@ -6823,11 +6232,10 @@

    Indonesian treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Indonesian treebanks. @@ -6839,35 +6247,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Irish + 3 + 168K + + + IE, Celtic + +
    -

    Irish treebanks

    -
    - - - -
    + +
    A Universal Dependencies 4910-sentence treebank for modern Irish. @@ -6889,12 +6291,11 @@

    Irish treebanks

  • Download
  • -

     

    - - - -
    + +
    + TwittIrish 47K @@ -6903,8 +6304,8 @@

    Irish treebanks

    -
    -
    + +
    A Universal Dependencies treebank of 2596 tweets in modern Irish. @@ -6916,12 +6317,11 @@

    Irish treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is the Cadhan Aonair UD treebank, consisting of 150 sentences randomly sampled from six pre-standard Irish texts. @@ -6947,11 +6347,10 @@

    Irish treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Irish treebanks. @@ -6963,35 +6362,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Italian + 11 + 1,020K + + + IE, Romance + +
    - -
    - - -

    Italian treebanks

    -
    - - - -
    + +
    The Italian corpus annotated according to the UD annotation scheme was obtained by conversion from ISDT (Italian Stanford Dependency Treebank), released for the dependency parsing shared task of Evalita-2014 (Bosco et al. 2014). @@ -7013,12 +6406,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + VIT 280K @@ -7027,8 +6419,8 @@

    Italian treebanks

    - -
    +
    +
    The UD_Italian-VIT corpus was obtained by conversion from VIT (Venice Italian Treebank), developed at the Laboratory of Computational Linguistics of the Università Ca' Foscari in Venice (Delmonte et al. 2007; Delmonte 2009; http://rondelmo.it/resource/VIT/Browser-VIT/index.htm). @@ -7040,12 +6432,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + KIParlaForest 18K @@ -7054,8 +6445,8 @@

    Italian treebanks

    - -
    +
    +
    The KIParla Forest treebank is a treebank of spoken Italian based on the [KIParla Corpus](https://kiparla.it/) @@ -7067,12 +6458,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + Valico 6K @@ -7081,8 +6471,8 @@

    Italian treebanks

    - -
    +
    +
    Manually corrected Treebank of Learner Italian drawn from the Valico corpus and correspondent corrected sentences. @@ -7094,12 +6484,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + Old 122K @@ -7108,8 +6497,8 @@

    Italian treebanks

    - -
    +
    +
    Italian-Old is a treebank containing **Dante Alighieri's Comedy** (composed between approximately 1306 and 1321), based on the 1994 Petrocchi edition and taken from the [**DanteSearch corpus**](https://dantesearch.dantenetwork.it), originally created at the University of Pisa, Italy. It is a treebank of Old Italian, specifically Florentine. @@ -7121,12 +6510,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + ParlaMint 20K @@ -7135,8 +6523,8 @@

    Italian treebanks

    - -
    +
    +
    ParlaMint-It is a collection of transcriptions of parliamentary sessions of the Italian Senate annotated in Universal Dependencies. The corpus is part of a larger multilingual collection of parliamentary transcripts built during the ParlaMint project (https://www.clarin.eu/parlamint). @@ -7148,12 +6536,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + ParTUT 55K @@ -7162,8 +6549,8 @@

    Italian treebanks

    - -
    +
    +
    UD_Italian-ParTUT is a conversion of a multilingual parallel treebank developed at the University of Turin, and consisting of a variety of text genres, including talks, legal texts and Wikipedia articles, among others. @@ -7176,12 +6563,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + TWITTIRO 29K @@ -7190,8 +6576,8 @@

    Italian treebanks

    - -
    +
    +
    TWITTIRÒ-UD is a collection of ironic Italian tweets annotated in Universal Dependencies. The treebank can be exploited for the training of NLP systems to enhance their performance on social media texts, and in particular, for irony detection purposes. @@ -7204,12 +6590,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + PoSTWITA 124K @@ -7218,8 +6603,8 @@

    Italian treebanks

    - -
    +
    +
    PoSTWITA-UD is a collection of Italian tweets annotated in Universal Dependencies that can be exploited for the training of NLP systems to enhance their performance on social media texts. @@ -7231,12 +6616,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 23K @@ -7245,8 +6629,8 @@

    Italian treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -7260,12 +6644,11 @@

    Italian treebanks

  • Download
  • -

     

    - - - -
    + +
    + MarkIT 40K @@ -7274,8 +6657,8 @@

    Italian treebanks

    - -
    +
    +
    The MarkIT resource contains around 800 sentences extracted from students' essays manually annotated with syntactic depencendies. The treebank covers seven types of marked constructions, plus some ambiguous sentences whose syntax can be wrongly classified as marked. @@ -7287,11 +6670,10 @@

    Italian treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Italian treebanks. @@ -7303,35 +6685,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Japanese + 6 + 2,645K + + + Japanese + +
    -

    Japanese treebanks

    -
    - - - -
    + +
    This Universal Dependencies (UD) Japanese treebank is based on the definition of UD Japanese convention described in the UD documentation. The original sentences are from Google UDT 2.0. @@ -7353,12 +6729,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - -
    + +
    + GSDLUW 150K @@ -7367,8 +6742,8 @@

    Japanese treebanks

    -
    -
    + +
    This Universal Dependencies (UD) Japanese treebank is based on the definition of UD Japanese convention described in the UD documentation. The original sentences are from Google UDT 2.0. @@ -7380,12 +6755,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -7409,12 +6783,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -7438,12 +6811,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - - -
    + +
    This Universal Dependencies (UD) Japanese treebank is based on the definition of UD Japanese convention described in the UD documentation. @@ -7467,12 +6839,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - - -
    + +
    This Universal Dependencies (UD) Japanese treebank is based on the definition of UD Japanese convention described in the UD documentation. @@ -7499,11 +6870,10 @@

    Japanese treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Japanese treebanks. @@ -7515,35 +6885,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Javanese + 1 + 14K + + + Austronesian, Javanese + +
    - -
    - - -

    Javanese treebanks

    -
    - - - -
    + +
    UD Javanese-CSUI is a dependency treebank in Javanese, a regional language in Indonesia with more than 68 million users. It was developed by Alfina et al. from the Faculty of Computer Science, Universitas Indonesia. The newest version has 1000 sentences and 14K words with manual annotation. @@ -7565,11 +6929,10 @@

    Javanese treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7579,35 +6942,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - - - -
    - +
    + + + Kaapor + 1 + <1K + + + Tupian, Maweti-Guarani + +
    -

    Kaapor treebanks

    -
    - - - -
    + +
    **UD_Kaapor-TuDeT** is a collection of annotated sentences in [Ka'apor](https://glottolog.org/resource/languoid/id/urub1250). The project is a work in progress and the treebank is being updated on a regular basis. @@ -7629,11 +6986,10 @@

    Kaapor treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7643,35 +6999,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Kadiweu + 1 + <1K + + + Guaicuruan + +
    - -
    - - -

    Kadiweu treebanks

    -
    - - - -
    + +
    UD_Kadiweu-UNICAMP is a treebank for [Kadiwéu](https://glottolog.org/resource/languoid/id/kadi1248) (ISO-639: `kbc`), an endangered Indigenous language of Brazil. It consists of isolated sentences produced by native speakers. @@ -7693,11 +7043,10 @@

    Kadiweu treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7707,35 +7056,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Kangri + 1 + 2K + + + IE, Indic + +
    -

    Kangri treebanks

    -
    - - - -
    + +
    The Kangri UD Treebank (KDTB) is a part of the Universal Dependency treebank project. @@ -7757,11 +7100,10 @@

    Kangri treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7771,35 +7113,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Karelian + 1 + 3K + + + Uralic, Finnic + +
    - -
    - - -

    Karelian treebanks

    -
    - - - -
    + +
    UD Karelian-KKPP is a manually annotated new corpus of Karelian made in Universal dependencies annotation scheme. The data is collected from @@ -7824,11 +7160,10 @@

    Karelian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7838,35 +7173,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Karo + 1 + 2K + + + Tupian, Ramarama + +
    -

    Karo treebanks

    -
    - - - -
    + +
    UD_Karo-TuDeT is a collection of annotated sentences in <a href="https://glottolog.org/resource/languoid/id/karo1306"> Karo</a>. The sentences stem from the only grammatical description of the language (Gabas, 1999) and from the sentences in the dictionary by the same author (Gabas, 2007). Sentence annotation and documentation by Fabrício Ferraz Gerardi. @@ -7888,11 +7217,10 @@

    Karo treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7902,35 +7230,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Kazakh + 1 + 10K + + + Turkic, Northwestern + +
    - -
    - - -

    Kazakh treebanks

    -
    - - - -
    + +
    The UD Kazakh treebank is a combination of text from various sources including Wikipedia, some folk tales, sentences from the UDHR, news and phrasebook sentences. Sentences IDs include partial document identifiers. @@ -7953,11 +7275,10 @@

    Kazakh treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -7967,35 +7288,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Khoekhoe + 1 + 29K + + + Khoe-Kwadi + +
    -

    Khoekhoe treebanks

    -
    - - - -
    + +
    UD\_Khoekhoe-KDT is a Universal Dependencies (UD) treebank for the Khoekhoegowab (Khoekhoe) language. The annotation was performed manually based on glosses. This treebank includes texts from various sources: fiction, grammar, and spoken conversation. The treebank contains **27k tokens**, distributed as follows: @@ -8021,11 +7336,10 @@

    Khoekhoe treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -8035,35 +7349,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Kiche + 1 + 10K + + + Mayan + +
    - -
    - - -

    Kiche treebanks

    -
    - - - -
    + +
    UD Kʼicheʼ-IU is a treebank consisting of sentences from a variety of text domains but principally dictionary example sentences and linguistic examples. @@ -8086,11 +7394,10 @@

    Kiche treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -8100,35 +7407,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Komi Permyak + 1 + 1K + + + Uralic, Permic + +
    -

    Komi Permyak treebanks

    -
    - - - -
    + +
    This is a Komi-Permyak literary language treebank consisting of original and translated texts. @@ -8150,11 +7451,10 @@

    Komi Permyak treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -8164,35 +7464,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Komi Zyrian + 2 + 10K + + + Uralic, Permic + +
    - -
    - - -

    Komi Zyrian treebanks

    -
    - - - -
    + +
    UD Komi-Zyrian Lattice is a treebank of written standard Komi-Zyrian. @@ -8214,12 +7508,11 @@

    Komi Zyrian treebanks

  • Download
  • -

     

    - - - -
    + +
    + IKDP 2K @@ -8228,8 +7521,8 @@

    Komi Zyrian treebanks

    - -
    +
    +
    This treebank consists of dialectal transcriptions of spoken Komi-Zyrian. The current texts are short recorded segments from different areas where the Iźva dialect of Komi language is spoken. @@ -8241,11 +7534,10 @@

    Komi Zyrian treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Komi Zyrian treebanks. @@ -8257,35 +7549,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Korean + 5 + 615K + + + Korean + +
    -

    Korean treebanks

    -
    - - - -
    + +
    UD_Korean-KSL is a dependency treebank of second-language (L2) Korean. @@ -8307,12 +7593,11 @@

    Korean treebanks

  • Download
  • -

     

    - - - -
    + +
    + Kaist 350K @@ -8321,8 +7606,8 @@

    Korean treebanks

    -
    -
    + +
    The KAIST Korean Universal Dependency Treebank is generated by Chun et al., 2018 from the constituency trees in the [KAIST Tree-Tagging Corpus](http://semanticweb.kaist.ac.kr/home/index.php/Corpus4). @@ -8334,12 +7619,11 @@

    Korean treebanks

  • Download
  • -

     

    - - - - -
    + +
    The Google Korean Universal Dependency Treebank is first converted from the [Universal Dependency Treebank v2.0 (legacy)](https://github.com/ryanmcd/uni-dep-tb), and then enhanced by Chun et al., 2018. @@ -8362,12 +7646,11 @@

    Korean treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -8391,12 +7674,11 @@

    Korean treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD Korean-LittlePrince is a UD adaptation of the k-SNACS dataset [(Hwang et al. 2020)](https://aclanthology.org/2020.dmr-1.6/). @@ -8419,11 +7701,10 @@

    Korean treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Korean treebanks. @@ -8435,35 +7716,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Kyrgyz + 2 + 25K + + + Turkic, Northwestern + +
    - -
    - - -

    Kyrgyz treebanks

    -
    - - - -
    + +
    UD_Kyrgyz-KTMU is dependency parsing based treebank in Kyrgyz language. The dataset mostly contains headlines from Kyrgyz news websites. @@ -8485,12 +7760,11 @@

    Kyrgyz treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL 1K @@ -8499,8 +7773,8 @@

    Kyrgyz treebanks

    -
    -
    + +
    This is a small treebank of grammatical examples for Kyrgyz. It is part of a parallel Universal Dependencies corpus containing 148 sentences across four Turkic languages, designed to facilitate cross-linguistic research on these related languages. @@ -8513,11 +7787,10 @@

    Kyrgyz treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Kyrgyz treebanks. @@ -8529,35 +7802,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Latgalian + 1 + <1K + + + IE, Baltic + +
    -

    Latgalian treebanks

    -
    - - - -
    + +
    UD_Latgalian-Cairo is an example treebank to provide minimal dataset for Latgalian based on the Cairo sample sentences. Created by [AI Lab](http://ailab.lv) at Institute of Mathematics and Computer Science, University of Latvia. @@ -8579,11 +7846,10 @@

    Latgalian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -8593,35 +7859,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Latin + 6 + 1,012K + + + IE, Italic + +
    - -
    - - -

    Latin treebanks

    -
    - - - -
    + +
    Latin data from the _Index Thomisticus_ Treebank. Data are taken from the _Index Thomisticus_ corpus by Roberto Busa SJ, which contains the complete work by Thomas Aquinas (1225–1274; Medieval Latin) and by 61 other authors related to Thomas. @@ -8643,12 +7903,11 @@

    Latin treebanks

  • Download
  • -

     

    - - - -
    + +
    + LLCT 242K @@ -8657,8 +7916,8 @@

    Latin treebanks

    - -
    +
    +
    This Universal Dependencies version of the **LLCT** (Late Latin Charter Treebank) consists of an automated conversion of the **LLCT2** treebank from the Latin Dependency Treebank (LDT) format into the Universal Dependencies standard. @@ -8670,12 +7929,11 @@

    Latin treebanks

  • Download
  • -

     

    - - - -
    + +
    + UDante 55K @@ -8684,8 +7942,8 @@

    Latin treebanks

    - -
    +
    +
    The **UDante** treebank is based on the Latin texts of Dante Alighieri, taken from the [**DanteSearch corpus**](https://dantesearch.dantenetwork.it), originally created at the University of Pisa, Italy. @@ -8699,12 +7957,11 @@

    Latin treebanks

  • Download
  • -

     

    - - - -
    + +
    + CIRCSE 29K @@ -8713,8 +7970,8 @@

    Latin treebanks

    - -
    +
    +
    UD_Latin-CIRCSE is a repository of treebanks featuring Latin texts natively annotated at the CIRCSE Research Centre in Milan (https://centridiricerca.unicatt.it/circse/en.html) following the Universal Dependencies (UD) (https://universaldependencies.org) annotation scheme. The repository includes prose and poetry texts from different periods. @@ -8727,12 +7984,11 @@

    Latin treebanks

  • Download
  • -

     

    - - - -
    + +
    + Perseus 29K @@ -8741,8 +7997,8 @@

    Latin treebanks

    - -
    +
    +
    This Universal Dependencies Latin Treebank consists of an automatic conversion of a selection of passages from the Ancient Greek and Latin @@ -8756,12 +8012,11 @@

    Latin treebanks

  • Download
  • -

     

    - - - -
    + +
    + PROIEL 205K @@ -8770,8 +8025,8 @@

    Latin treebanks

    - -
    +
    +
    The Latin PROIEL treebank is based on the Latin data from the PROIEL treebank, and contains most of the Vulgate New Testament translations plus selections from Caesar's Gallic War, Cicero's Letters to Atticus, Palladius' Opus Agriculturae and the first book of Cicero's De officiis. @@ -8783,11 +8038,10 @@

    Latin treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Latin treebanks. @@ -8799,35 +8053,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Latvian + 2 + 330K + + + IE, Baltic + +
    -

    Latvian treebanks

    -
    - - - -
    + +
    Latvian UD Treebank is based on Latvian Treebank ([LVTB](http://sintakse.korpuss.lv)), being created at University of Latvia, Institute of Mathematics and Computer Science, [Artificial Intelligence Laboratory](http://ailab.lv). @@ -8849,12 +8097,11 @@

    Latvian treebanks

  • Download
  • -

     

    - - - -
    + +
    + Cairo <1K @@ -8863,8 +8110,8 @@

    Latvian treebanks

    -
    -
    + +
    This is an example treebank made to ilustrate UD annotation choices made for Latvian based on the Cairo sample sentences. Created by [AI Lab](http://ailab.lv) at Institute of Mathematics and Computer Science, University of Latvia. @@ -8876,11 +8123,10 @@

    Latvian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Latvian treebanks. @@ -8892,35 +8138,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Ligurian + 1 + 6K + + + IE, Romance + +
    - -
    - - -

    Ligurian treebanks

    -
    - - - -
    + +
    The Genoese Ligurian Treebank is a small, manually annotated collection of contemporary Ligurian prose. The focus of the treebank is written Genoese, the koiné variety of Ligurian which is associated with today's literary, journalistic and academic ligurophone sphere. @@ -8942,11 +8182,10 @@

    Ligurian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -8956,35 +8195,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Lithuanian + 2 + 75K + + + IE, Baltic + +
    -

    Lithuanian treebanks

    -
    - - - -
    + +
    The Lithuanian dependency treebank ALKSNIS v3.0 (Vytautas Magnus University). @@ -9006,12 +8239,11 @@

    Lithuanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + HSE 5K @@ -9020,8 +8252,8 @@

    Lithuanian treebanks

    -
    -
    + +
    Lithuanian treebank annotated manually (dependencies) using the Morphological Annotator by CCL, Vytautas Magnus University (http://tekstynas.vdu.lt/) and manual disambiguation. A pilot version which includes news and an essay by Tomas Venclova is available here. @@ -9034,11 +8266,10 @@

    Lithuanian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Lithuanian treebanks. @@ -9050,35 +8281,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Livvi + 1 + 1K + + + Uralic, Finnic + +
    - -
    - - -

    Livvi treebanks

    -
    - - - -
    + +
    UD Livvi-KKPP is a manually annotated new corpus of Livvi-Karelian made directly in the Universal dependencies annotation scheme. The data is collected from @@ -9103,11 +8328,10 @@

    Livvi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9117,35 +8341,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Low Saxon + 1 + 22K + + + IE, Germanic + +
    -

    Low Saxon treebanks

    -
    - - - -
    + +
    The UD Low Saxon LSDC dataset consists of sentences in 8 major Low Saxon dialect groups from both Germany and the Netherlands. These sentences are (or are to become) part of the LSDC dataset and represent the language from mostly the 19th and early 20th century in genres such as short stories, novels, speeches, letters and fairytales. @@ -9167,11 +8385,10 @@

    Low Saxon treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9181,35 +8398,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Luxembourgish + 1 + <1K + + + IE, Germanic + +
    - -
    - - -

    Luxembourgish treebanks

    -
    - - - -
    + +
    The LuxBank corpus currently consists of the translated Cairo Cicling examples, and will be extended to include examples from a national dataset. It is the first comprehensive tree bank dataset for Luxembourgish. @@ -9231,11 +8442,10 @@

    Luxembourgish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9245,35 +8455,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Macedonian + 1 + 1K + + + IE, Slavic + +
    -

    Macedonian treebanks

    -
    - - - -
    + +
    The Macedonian-MTB treebank is a collection of annotated sentences taken from the Macedonian version of the Cairo CICLing Corpus and from the university textbook in syntax "Contemporary Macedonian Language 4" by Simov Sazdov. @@ -9295,11 +8499,10 @@

    Macedonian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9309,35 +8512,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Madi + 1 + <1K + + + Arawan + +
    - -
    - - -

    Madi treebanks

    -
    - - - -
    + +
    UD_Madi-Jarawara is a collection of annotated sentences in Madí (Jarawara dialect) from a variety of sources, including grammar examples, oral stories, didatic material, and dictionary examples. @@ -9360,11 +8557,10 @@

    Madi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9374,35 +8570,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Maghrebi Arabic French + 1 + 19K + + + Code switching + +
    -

    Maghrebi Arabic French treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for a romanized user-generated content variety of Algerian, a North-African Arabic dialect known for its frequent usage of code-switching. We added to the UD annotations NER annotations extending the French Treebank NER scheme (Sagot et al, 2012) and Offensive language classification and corrected many of the translations (still ongoing). @@ -9424,11 +8614,10 @@

    Maghrebi Arabic French treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9438,35 +8627,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Makurap + 1 + <1K + + + Tupian, Tupari + +
    - -
    - - -

    Makurap treebanks

    -
    - - - -
    + +
    UD_Makuráp-TuDeT is a collection of annotated texts in Makuráp. The project is a work in progress and the treebank is being updated on a regular basis. The sentences are being annotated by Carolina Aragon, Fabrício Ferraz Gerardi, Luana dos Santos, and Luan Cabral. @@ -9488,11 +8671,10 @@

    Makurap treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9502,35 +8684,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Malayalam + 1 + 2K + + + Dravidian + +
    -

    Malayalam treebanks

    -
    - - - -
    + +
    Currently just a small sample of Malayalam grammatical examples. @@ -9552,11 +8728,10 @@

    Malayalam treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9566,35 +8741,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Maltese + 1 + 44K + + + Afro-Asiatic, Semitic + +
    - -
    - - -

    Maltese treebanks

    -
    - - - -
    + +
    MUDT (Maltese Universal Dependencies Treebank) is a manually annotated treebank of Maltese, a Semitic language of Malta descended from North African Arabic with a significant amount of Italo-Romance influence. MUDT was designed as a balanced corpus with four major genres (see Splitting below) represented roughly equally. @@ -9617,11 +8786,10 @@

    Maltese treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9631,35 +8799,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Manx + 1 + 20K + + + IE, Celtic + +
    -

    Manx treebanks

    -
    - - - -
    + +
    This is the Cadhan Aonair UD treebank for Manx Gaelic, created by Kevin Scannell. @@ -9682,11 +8844,10 @@

    Manx treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9696,35 +8857,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Marathi + 2 + 122K + + + IE, Indic + +
    - -
    - - -

    Marathi treebanks

    -
    - - - -
    + +
    UD Marathi is a manually annotated treebank consisting primarily of stories from Wikisource, and parts of an article on Wikipedia. @@ -9746,12 +8901,11 @@

    Marathi treebanks

  • Download
  • -

     

    - - - -
    + +
    + CMUPAN 118K @@ -9760,8 +8914,8 @@

    Marathi treebanks

    - -
    +
    +
    This treebank is a modified version of a semi-automatically treebank authord by Aditi Chaudhary, which in turn is based on the treebanks released by KCIS, IIIT-Hyderabad. Additionally, the treebank also contains Marathi-Discourse: A manually annotated 35-sentence corpus covering political discourse. @@ -9773,11 +8927,10 @@

    Marathi treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Marathi treebanks. @@ -9789,35 +8942,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Mbya Guarani + 1 + 1K + + + Tupian, Maweti-Guarani + +
    -

    Mbya Guarani treebanks

    -
    - - - -
    + +
    UD Mbya_Guarani-Thomas is a corpus of Mbyá Guaraní (Tupian) texts collected by Guillaume Thomas. The current version of the corpus consists of three speeches by Paulina Kerechu Núñez Romero, a Mbyá Guaraní speaker from Ytu, Caazapá Department, Paraguay. @@ -9839,11 +8986,10 @@

    Mbya Guarani treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Mbya Guarani treebanks. @@ -9855,35 +9001,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Middle Armenian + 1 + 1K + + + IE, Armenian + +
    - -
    - - -

    Middle Armenian treebanks

    -
    - - - -
    + +
    A Universal Dependencies treebank for Middle Armenian developed for UD originally by the ArmTDP team led by Marat M. Yavrumyan at the Yerevan State University. @@ -9905,11 +9045,10 @@

    Middle Armenian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -9919,35 +9058,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Middle French + 2 + 126K + + + IE, Romance + +
    -

    Middle French treebanks

    -
    - - - -
    + +
    Middle-French ALTM (AUTOMATED Legal Texts Medieval) is a treebank of medieval legal French from Normandy. Currently in contains one text, an extract from _Coutume, style et usage au temps des Échiquiers de Normandie_, dated 1425. @@ -9970,12 +9103,11 @@

    Middle French treebanks

  • Download
  • -

     

    - - - -
    + +
    + PROFITEROLE 119K @@ -9984,8 +9116,8 @@

    Middle French treebanks

    -
    -
    + +
    UD_Middle_French-PROFITEROLE is the Middle French section of the PROFITEROLE corpus, the Old French section is UD_OLD_FRENCH-PROFITEROLE. @@ -9997,11 +9129,10 @@

    Middle French treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Middle French treebanks. @@ -10013,35 +9144,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Moksha + 1 + 4K + + + Uralic, Mordvin + +
    - -
    - - -

    Moksha treebanks

    -
    - - - -
    + +
    Erme Universal Dependencies annotated texts Moksha are the origin of UD_Moksha-JR with annotation (CoNLL-U) for texts in the Moksha language, it originally consists of a sample from a number of fiction authors writing originals in Moksha. @@ -10064,11 +9189,10 @@

    Moksha treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10078,35 +9202,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Munduruku + 1 + 1K + + + Tupian, Munduruku + +
    -

    Munduruku treebanks

    -
    - - - -
    + +
    UD_Munduruku-TuDeT is a collection of annotated sentences in [Mundurukú](http://www.endangeredlanguages.com/lang/2981). The project is a work in progress and the treebank is being updated on a regular basis. @@ -10134,11 +9252,10 @@

    Munduruku treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10148,35 +9265,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Naga + 1 + 3K + + + Sino-Tibetan, Tangkhul-Maring + +
    - -
    - - -

    Naga treebanks

    -
    - - - -
    + +
    UD_Naga-Suansu is a Universal Dependencies (UD) treebank for Suansu (Glottocode: suan1234), an endangered Tibeto-Burman language spoken on the Indo-Myanmar border. The annotation was performed manually based on glosses. This treebank includes texts from fiction and grammar. The treebank contains **3.1k** tokens, distributed as follows: @@ -10201,11 +9312,10 @@

    Naga treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10215,35 +9325,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Naija + 1 + 140K + + + Creole + +
    -

    Naija treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for spoken Naija (Nigerian Pidgin). @@ -10265,11 +9369,10 @@

    Naija treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10279,35 +9382,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Neapolitan + 1 + <1K + + + IE, Romance + +
    - -
    - - -

    Neapolitan treebanks

    -
    - - - -
    + +
    This treebank contains example sentences in Neapolitan, translated by a native speaker. @@ -10329,11 +9426,10 @@

    Neapolitan treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10343,35 +9439,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Nenets + 1 + 1K + + + Uralic, Samoyedic + +
    -

    Nenets treebanks

    -
    - - - -
    + +
    The Tundra Nenets UD treebank is converted from the [Tundra Nenets mSUD treebank](https://github.com/surfacesyntacticud/mSUD_Nenets-Tundra). The conversion from mSUD to UD is performed automatically followed by a comprehensive manual revision to ensure compliance with the UD annotation standards. @@ -10393,11 +9483,10 @@

    Nenets treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10407,35 +9496,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Nepali + 1 + <1K + + + IE, Indic + +
    - -
    - - -

    Nepali treebanks

    -
    - - - -
    + +
    UD_Nepali-BK is a manually annotated Universal Dependencies treebank for Nepali, an Indo-Aryan language written in Devanagari. The treebank contains sentences from a fictional narrative story and an argumentative discourse text, and follows the Universal Dependencies v2 guidelines. @@ -10457,11 +9540,10 @@

    Nepali treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10471,35 +9553,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Nheengatu + 1 + 26K + + + Tupian, Maweti-Guarani + +
    -

    Nheengatu treebanks

    -
    - - - -
    + +
    [UD_Nheengatu-CompLin](https://aclanthology.org/2024.propor-2.8) is a treebank of [Nheengatu](https://glottolog.org/resource/languoid/id/nhen1239), also known as Modern Tupi and *Língua Geral Amazônica* (ISO 639: `yrl`). It comprises sentences drawn from a wide range of published sources, including spontaneous speech, grammatical descriptions, fables, myths, coursebooks, and dictionaries. @@ -10521,11 +9597,10 @@

    Nheengatu treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10535,35 +9610,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + North Sami + 1 + 26K + + + Uralic, Sami + +
    - -
    - - -

    North Sami treebanks

    -
    - - - -
    + +
    This is a North Sámi treebank based on a manually disambiguated and function-labelled gold-standard corpus of North Sámi produced by the Giellatekno team at UiT Norgga árktalaš universitehta. @@ -10586,11 +9655,10 @@

    North Sami treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10600,35 +9668,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Northern Kurdish + 1 + 10K + + + IE, Iranian + +
    -

    Northern Kurdish treebanks

    -
    - - - -
    + +
    The treebank is a corpus of Kurmanji Kurdish. It contains fiction and encyclopaedic texts in roughly equal measure. It has been annotated natively in accordance with the UD annotation scheme. @@ -10651,11 +9713,10 @@

    Northern Kurdish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10665,35 +9726,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Northwest Gbaya + 1 + 2K + + + Niger-Congo, Gbaya-Manza-Ngbaka + +
    - -
    - - -

    Northwest Gbaya treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for Northwest Gbaya, a member of the Gbaya branch of the Atlantic-Congo phylum. The language is mainly spoken by about 250,000 speakers in Central African Republic. @@ -10715,11 +9770,10 @@

    Northwest Gbaya treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10729,35 +9783,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Norwegian + 2 + 611K + + + IE, Germanic + +
    -

    Norwegian treebanks

    -
    - - - -
    + +
    The Norwegian UD treebank is based on the Bokmål section of the Norwegian Dependency Treebank (NDT), which is a syntactic treebank of Norwegian. The current version of NDT has been automatically converted to the UD @@ -10782,12 +9830,11 @@

    Norwegian treebanks

  • Download
  • -

     

    - - - -
    + +
    + Nynorsk 301K @@ -10796,8 +9843,8 @@

    Norwegian treebanks

    -
    -
    + +
    The Norwegian UD treebank is based on the Nynorsk section of the Norwegian Dependency Treebank (NDT), which is a syntactic treebank of Norwegian. @@ -10812,11 +9859,10 @@

    Norwegian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Norwegian treebanks. @@ -10828,35 +9874,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Occitan + 1 + 25K + + + IE, Romance + +
    - -
    - - -

    Occitan treebanks

    -
    - - - -
    + +
    Tolosa Treebank was developed as part of the EFA 227/16 LINGUATEC Project, financed by the POCTEFA Interreg European funds. It includes data from literature, newspapers, encyclopedia, scientific papers and web blogs. @@ -10879,11 +9919,10 @@

    Occitan treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10893,35 +9932,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Odia + 1 + 5K + + + IE, Indic + +
    -

    Odia treebanks

    -
    - - - -
    + +
    The Odia UD Treebank (ODTB) is a part of the Universal Dependency treebank project. @@ -10943,11 +9976,10 @@

    Odia treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -10957,35 +9989,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Old Church Slavonic + 1 + 198K + + + IE, Slavic + +
    - -
    - - -

    Old Church Slavonic treebanks

    -
    - - - -
    + +
    The Old Church Slavonic (OCS) UD treebank is based on canonical Old Church Slavonic data from the PROIEL and TOROT treebanks. @@ -11007,11 +10033,10 @@

    Old Church Slavonic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -11021,35 +10046,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old East Slavic + 4 + 580K + + + IE, Slavic + +
    -

    Old East Slavic treebanks

    -
    - - - -
    + +
    `UD_Old_East_Slavic-RNC` is a sample of the Middle Russian corpus (1300-1700), a part of the Russian National Corpus. The data were originally annotated according to the RNC and extended UD-Russian morphological schemas and UD 2.4 dependency schema. @@ -11071,12 +10090,11 @@

    Old East Slavic treebanks

  • Download
  • -

     

    - - - -
    + +
    + Ruthenian 137K @@ -11085,8 +10103,8 @@

    Old East Slavic treebanks

    -
    -
    + +
    The Ruthenian UD treebank includes texts written in the territories of modern Belarus, Lithuania, Ukraine, and Poland in ca. 1300-1700. A sample of legal and nonfiction texts is drawn from the Ruthenian Corpus. @@ -11098,12 +10116,11 @@

    Old East Slavic treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD\_Old\_East\_Slavic-TOROT is a conversion of a selection of Old East Slavonic and Middle Russian data from the Tromsø Old Russian and OCS Treebank (TOROT), which was originally annotated in PROIEL dependency format. @@ -11125,12 +10142,11 @@

    Old East Slavic treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD Old\_East\_Slavic-Birchbark is based on the RNC Corpus of Birchbark Letters and includes documents written in 1025-1500 in an East Slavic vernacular (letters, household and business records, @@ -11156,11 +10172,10 @@

    Old East Slavic treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Old East Slavic treebanks. @@ -11172,35 +10187,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Old English + 1 + <1K + + + IE, Germanic + +
    - -
    - - -

    Old English treebanks

    -
    - - - -
    + +
    Old English [Cairo](https://github.com/UniversalDependencies/cairo) sentences with UD and additional annotations @@ -11222,11 +10231,10 @@

    Old English treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -11236,35 +10244,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old French + 2 + 255K + + + IE, Romance + +
    -

    Old French treebanks

    -
    - - - -
    + +
    Old-French ALTM (AUTOMATED Legal Texts Medieval) is a treebank of medieval legal French from Normandy. Currently in contains one text, _Atiremens et jugiés d'eschequiers_, dated 1314. @@ -11287,12 +10289,11 @@

    Old French treebanks

  • Download
  • -

     

    - - - -
    + +
    + PROFITEROLE 240K @@ -11301,8 +10302,8 @@

    Old French treebanks

    -
    -
    + +
    UD_Old_French-PROFITEROLE is an expansion of the previous UD_Old_French-SRCMF (which was a conversion of (part of) the SRCMF corpus (Syntactic Reference Corpus of Medieval French @@ -11316,11 +10317,10 @@

    Old French treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Old French treebanks. @@ -11332,35 +10332,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Old Georgian + 1 + 6K + + + Kartvelian + +
    - -
    - - -

    Old Georgian treebanks

    -
    - - - -
    + +
    The Old Georgian UD Treebank (UD_Old_Georgian-GLC) is the first syntactically annotated corpus of Georgian, based on a collection of annotated sentences selected from the Old Georgian Language Corpus (OGLC) available at https://oge.iliauni.edu.ge/. @@ -11382,11 +10376,10 @@

    Old Georgian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -11396,35 +10389,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old Irish + 2 + <1K + + + IE, Celtic + +
    -

    Old Irish treebanks

    -
    - - - -
    + +
    A Universal Dependencies treebank for the Old Irish Würzburg glosses. @@ -11446,12 +10433,11 @@

    Old Irish treebanks

  • Download
  • -

     

    - - - -
    + +
    + DipSGG <1K @@ -11460,8 +10446,8 @@

    Old Irish treebanks

    -
    -
    + +
    A Universal Dependencies treebank for the Old Irish glosses of St. Gall. @@ -11473,11 +10459,10 @@

    Old Irish treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Old Irish treebanks. @@ -11489,35 +10474,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Old Occitan + 1 + 52K + + + IE, Romance + +
    - -
    - - -

    Old Occitan treebanks

    -
    - - - -
    + +
    UD_Old_Occitan-CorAG (Corpus de l'Ancien Gascon) is a corpus of medieval and early modern legal texts in Gascon, a variety of Old Occitan. The texts were digitized from existing editions and subsequently manually annotated in Universal Dependencies (PoS, functions and some morphological features). @@ -11539,11 +10518,10 @@

    Old Occitan treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -11553,35 +10531,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old Turkish + 1 + <1K + + + Turkic, Northeastern + +
    -

    Old Turkish treebanks

    -
    - - - -
    + +
    This repository contains an [Old Turkish](https://iso639-3.sil.org/code/otk) treebank built upon Old Turkic script texts. @@ -11603,11 +10575,10 @@

    Old Turkish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -11617,35 +10588,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Ottoman Turkish + 3 + 31K + + + Turkic, Southwestern + +
    - -
    - - -

    Ottoman Turkish treebanks

    -
    - - - -
    + +
    An Ottoman Turkish dependency treebank annotated in UD style. Created by Enes Yılandiloğlu. @@ -11667,12 +10632,11 @@

    Ottoman Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL <1K @@ -11681,8 +10645,8 @@

    Ottoman Turkish treebanks

    - -
    +
    +
    The Ottoman Turkish-TueCL treebank is part of a parallel Universal Dependencies corpus containing 148 sentences across five Turkic languages (Turkish, Azerbaijani, Kyrgyz, Uzbek, and Ottoman Turkish), designed to facilitate cross-linguistic research on these related languages. @@ -11694,12 +10658,11 @@

    Ottoman Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + BOUN 8K @@ -11708,8 +10671,8 @@

    Ottoman Turkish treebanks

    - -
    +
    +
    An Ottoman Turkish dependency treebank annotated in UD style. Created by [Şaziye Betül Özateş](https://sb-b.github.io/), Tarık Emre Tıraş, Efe Eren Genç from Boğaziçi University, and Esma Fatıma Bilgin Taşdemir from Medeniyet University. @@ -11721,11 +10684,10 @@

    Ottoman Turkish treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Ottoman Turkish treebanks. @@ -11737,35 +10699,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Pashto + 2 + 6K + + + IE, Iranian + +
    -

    Pashto treebanks

    -
    - - - -
    + +
    The Pashto-Sikaram treebank is a native UD treebank with manually annotated texts from various sources. @@ -11787,12 +10743,11 @@

    Pashto treebanks

  • Download
  • -

     

    - - - -
    + +
    + Prince 1K @@ -11801,8 +10756,8 @@

    Pashto treebanks

    -
    -
    + +
    The UD Pashto-Prince treebank contains manually annotated Pashto sentences from two textual sources: 50 sentences from Le Petit Prince, which was then translated and adapted into Northern Pashto, and 14 sentences from a Pashto prose text on Pashtun leadership. All sentences are annotated natively according to Universal Dependencies guidelines. @@ -11814,11 +10769,10 @@

    Pashto treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Pashto treebanks. @@ -11830,35 +10784,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Paumari + 1 + <1K + + + Arawan + +
    - -
    - - -

    Paumari treebanks

    -
    - - - -
    + +
    This is a small treebank of Paumari, a low-resource Amazonian language. @@ -11880,11 +10828,10 @@

    Paumari treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -11894,35 +10841,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Persian + 2 + 654K + + + IE, Iranian + +
    -

    Persian treebanks

    -
    - - - -
    + +
    The Persian Universal Dependency Treebank (PerUDT) is the result of automatic coversion of Persian Dependency Treebank (PerDT) with extensive manual corrections. Please refer to the follwoing work, if you use this data: * Mohammad Sadegh Rasooli, Pegah Safari, Amirsaeid Moloodi, and Alireza Nourian. "The Persian Dependency Treebank Made Universal". 2020 (to appear). @@ -11945,12 +10886,11 @@

    Persian treebanks

  • Download
  • -

     

    - - - -
    + +
    + Seraji 152K @@ -11959,8 +10899,8 @@

    Persian treebanks

    -
    -
    + +
    The Persian Universal Dependency Treebank (Seraji) is based on Uppsala Persian Dependency Treebank (UPDT). The conversion of the UPDT to the Universal Dependencies was performed semi-automatically with extensive manual checks and corrections. @@ -11973,11 +10913,10 @@

    Persian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Persian treebanks. @@ -11989,35 +10928,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Pesh + 1 + 4K + + + Chibchan, Pesh + +
    - -
    - - -

    Pesh treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for Pesh (aka Paya), a member of the Chibchan language family. The language is spoken by about 500 speakers in Honduras. @@ -12039,11 +10972,10 @@

    Pesh treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -12053,35 +10985,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Phrygian + 1 + 1K + + + IE, Greek + +
    -

    Phrygian treebanks

    -
    - - - -
    + +
    UD Phrygian-KUL started as part of a Master's thesis in linguistics at KU Leuven, annotating the New Phrygian subcorpus of the ancient Phrygian language. It has since expanded to include Old and Middle Phrygian texts. @@ -12103,11 +11029,10 @@

    Phrygian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -12117,35 +11042,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Polish + 4 + 546K + + + IE, Slavic + +
    - -
    - - -

    Polish treebanks

    -
    - - - -
    + +
    The Polish PDB-UD treebank is automatically converted from the Polish Dependency Bank 2.0 (PDB 2.0). Both treebanks were created at the [Institute of Computer Science, Polish Academy of Sciences](https://ipipan.waw.pl/en/) in Warsaw (Poland). @@ -12167,12 +11086,11 @@

    Polish treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 18K @@ -12181,8 +11099,8 @@

    Polish treebanks

    - -
    +
    +
    This is the Polish portion of the Parallel Universal Dependencies (PUD) treebanks, created at the [Institute of Computer Science, Polish Academy of Sciences](https://ipipan.waw.pl/en/) in Warsaw (Poland). @@ -12194,12 +11112,11 @@

    Polish treebanks

  • Download
  • -

     

    - - - -
    + +
    + LFG 130K @@ -12208,8 +11125,8 @@

    Polish treebanks

    - -
    +
    +
    The LFG Enhanced UD treebank of Polish is based on a corpus of LFG (Lexical Functional Grammar) syntactic structures generated by an LFG grammar of Polish, POLFIE, and manually disambiguated by human annotators. @@ -12221,12 +11138,11 @@

    Polish treebanks

  • Download
  • -

     

    - - - -
    + +
    + MPDT 47K @@ -12235,8 +11151,8 @@

    Polish treebanks

    - -
    +
    +
    UD_Polish-MPDT is a treebank of Middle Polish (17th–18th centuries). It is a rule-based conversion of the [Middle Polish Dependency Treebank](https://korba.edu.pl/treebank?lang=en) (Wieczorek, 2025) from its original annotation to the Universal Dependencies format. The MPDT sentences are sourced from the [KorBa corpus](https://korba.edu.pl/overview?lang=en) (Gruszczyński et al., 2022). @@ -12248,11 +11164,10 @@

    Polish treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Polish treebanks. @@ -12264,35 +11179,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Pomak + 1 + 34K + + + IE, Slavic + +
    -

    Pomak treebanks

    -
    - - - -
    + +
    The Pomak UD treebank is derived from the Pomak Dependency Treebank, a resource developed and maintained by researchers at the Institute for Language and Speech Processing/Athena @@ -12316,11 +11225,10 @@

    Pomak treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -12330,35 +11238,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Portuguese + 7 + 1,545K + + + IE, Romance + +
    - -
    - - -

    Portuguese treebanks

    -
    - - - -
    + +
    Porttinari-base [(Duran et al., 2023)](https://sol.sbc.org.br/index.php/stil/article/view/25443/25264) is the journalistic portion of Porttinari (which stands for “PORTuguese Treebank”), which shall be a large multigenre treebank for Portuguese [(Pardo et al., 2021)](https://sol.sbc.org.br/index.php/stil/article/view/17778/17612), following the "Universal Dependencies" international grammar framework [(de Marneffe et al., 2021)](https://aclanthology.org/2021.cl-2.11/). @@ -12380,12 +11282,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + DANTEStocks 80K @@ -12394,8 +11295,8 @@

    Portuguese treebanks

    - -
    +
    +
    DANTEStocks (Di Felippo et al., 2024) is a collection of Brazilian Portuguese tweets on the stock market domain that is part of Porttinari (“PORTuguese Treebank”), which shall be a large multigenre treebank for Portuguese (Pardo et al., 2021), following the "Universal Dependencies" framework (de Marneffe et al., 2021). @@ -12407,12 +11308,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + PetroGold 250K @@ -12421,8 +11321,8 @@

    Portuguese treebanks

    - -
    +
    +
    UD_Portuguese-PetroGold is a fully revised treebank which consists of academic texts from the oil & gas domain in Brazilian Portuguese. @@ -12434,12 +11334,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + Bosque 227K @@ -12448,8 +11347,8 @@

    Portuguese treebanks

    - -
    +
    +
    This Universal Dependencies (UD) Portuguese treebank is based on the Constraint Grammar converted version of the Bosque, which is part of @@ -12464,12 +11363,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + CINTIL 475K @@ -12478,8 +11376,8 @@

    Portuguese treebanks

    - -
    +
    +
    CINTIL-UDep is a dependency bank of Portuguese that is treebanked with Universal Dependencies. It contains over 38K annotated sentences (and 476K tokens), of mostly newspaper text. @@ -12492,12 +11390,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + GSD 318K @@ -12506,8 +11403,8 @@

    Portuguese treebanks

    - -
    +
    +
    The Brazilian Portuguese UD is converted from the [Google Universal Dependency Treebank v2.0 (legacy)](https://github.com/ryanmcd/uni-dep-tb). @@ -12520,12 +11417,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 23K @@ -12534,8 +11430,8 @@

    Portuguese treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -12549,11 +11445,10 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Portuguese treebanks. @@ -12565,35 +11460,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Punjabi + 2 + 3K + + + IE, Indic + +
    -

    Punjabi treebanks

    -
    - - - -
    + +
    The UD_Punjabi_CS is a manually annotated treebank in Punjabi (also called Eastern Punjabi, Gurmukhi script) language. The Indo-Aryan language is prominently spoken in Punjab region of India and Pakistan with strong diasporas in Western countries like Canada, UK and USA etc. It is written from Left-to-Right with Subject-Object-Verb (SOV) word ordering. @@ -12615,12 +11504,11 @@

    Punjabi treebanks

  • Download
  • -

     

    - - - -
    + +
    + Rang 1K @@ -12629,8 +11517,8 @@

    Punjabi treebanks

    -
    -
    + +
    The Punjabi-Rang treebank is a manually annotated corpus in Punjabi (Shahmukhi script). @@ -12642,11 +11530,10 @@

    Punjabi treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Punjabi treebanks. @@ -12658,35 +11545,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Romanian + 6 + 942K + + + IE, Romance + +
    - -
    - - -

    Romanian treebanks

    -
    - - - -
    + +
    The UD treebank ArT is a treebank of the Aromanian dialect of the Romanian language in UD format. @@ -12708,12 +11589,11 @@

    Romanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + MolDoRo <1K @@ -12722,8 +11602,8 @@

    Romanian treebanks

    - -
    +
    +
    A small treebank of sentences in Moldovan Romanian, using the Cyrillic writing system (as used in Moldova until 1989). @@ -12735,12 +11615,11 @@

    Romanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + Nonstandard 572K @@ -12749,8 +11628,8 @@

    Romanian treebanks

    - -
    +
    +
    The Romanian Non-standard UD treebank (called UAIC-RoDia) is based on UAIC-RoDia Treebank. UAIC-RoDia = ISLRN 156-635-615-024-0 @@ -12762,12 +11641,11 @@

    Romanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + RRT 218K @@ -12776,8 +11654,8 @@

    Romanian treebanks

    - -
    +
    +
    The Romanian UD treebank (called RoRefTrees) (Barbu Mititelu et al., 2016) is the reference treebank in UD format for standard Romanian. @@ -12789,12 +11667,11 @@

    Romanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + SiMoNERo 146K @@ -12803,8 +11680,8 @@

    Romanian treebanks

    - -
    +
    +
    SiMoNERo is a medical corpus of contemporary Romanian. @@ -12816,12 +11693,11 @@

    Romanian treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL 4K @@ -12830,8 +11706,8 @@

    Romanian treebanks

    - -
    +
    +
    The Romanian Social Media Sexist Language UD Treebank is a reference treebank in Universal Dependencies (UD) format for Romanian sexist language. Currently small, it comprises a subset of tweets sourced from [CoRoSeOf](https://github.com/DianaHoefels/CoRoSeOf). @@ -12843,11 +11719,10 @@

    Romanian treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Romanian treebanks. @@ -12859,35 +11734,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Russian + 5 + 3,455K + + + IE, Slavic + +
    -

    Russian treebanks

    -
    - - - -
    + +
    Universal Dependencies treebank is based on data samples extracted from Taiga Corpus and MorphoRuEval-2017 and GramEval-2020 shared tasks collections. @@ -12909,12 +11778,11 @@

    Russian treebanks

  • Download
  • -

     

    - - - -
    + +
    + SynTagRus 1,515K @@ -12923,8 +11791,8 @@

    Russian treebanks

    -
    -
    + +
    Russian data from the SynTagRus corpus. @@ -12936,12 +11804,11 @@

    Russian treebanks

  • Download
  • -

     

    - - - - -
    + +
    Russian Universal Dependencies Treebank annotated and converted by Google. @@ -12963,12 +11830,11 @@

    Russian treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD_Russian-Poetry contains samples of Russian poetry written in 19th – early 21th centuries. The treebank is based on the Poetry Corpus of the Russian National Corpus. @@ -12990,12 +11856,11 @@

    Russian treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -13019,11 +11884,10 @@

    Russian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Russian treebanks. @@ -13035,35 +11899,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Ruuli + 1 + 6K + + + Niger-Congo, Bantoid + +
    - -
    - - -

    Ruuli treebanks

    -
    - - - -
    + +
    UD_Ruuli-RDT is a Universal Dependencies (UD) treebank for the Ruruuli-Lunyala (Ruuli) language. The annotation was converted from interlinear glossed text and manually annotated for syntactic relations. The treebank includes texts from various sources: conversations, oral folktales, biographic monologue, movie subtitles, grammar examples, and factual prose. The treebank contains approximately 6,000 tokens. @@ -13085,11 +11943,10 @@

    Ruuli treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13099,35 +11956,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Sanskrit + 2 + 208K + + + IE, Indic + +
    -

    Sanskrit treebanks

    -
    - - - -
    + +
    A small Sanskrit treebank of sentences from Pañcatantra, an ancient Indian collection of interrelated fables by Vishnu Sharma. @@ -13150,12 +12001,11 @@

    Sanskrit treebanks

  • Download
  • -

     

    - - - -
    + +
    + Vedic 206K @@ -13164,8 +12014,8 @@

    Sanskrit treebanks

    -
    -
    + +
    The Treebank of Vedic Sanskrit contains 4,000 sentences with 27,000 words chosen from metrical and prose passages of the Ṛgveda (RV), the Śaunaka recension of the Atharvaveda (ŚS), the Maitrāyaṇīsaṃhitā (MS), and the Aitareya- (AB) and Śatapatha-Brāhmaṇas (ŚB). @@ -13180,11 +12030,10 @@

    Sanskrit treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Sanskrit treebanks. @@ -13196,35 +12045,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Scottish Gaelic + 1 + 90K + + + IE, Celtic + +
    - -
    - - -

    Scottish Gaelic treebanks

    -
    - - - -
    + +
    A treebank of Scottish Gaelic based on the [Annotated Reference Corpus Of Scottish Gaelic (ARCOSG)](https://github.com/Gaelic-Algorithmic-Research-Group/ARCOSG). @@ -13247,11 +12090,10 @@

    Scottish Gaelic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13261,35 +12103,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Serbian + 1 + 97K + + + IE, Slavic + +
    -

    Serbian treebanks

    -
    - - - -
    + +
    The Serbian UD treebank is based on the [SETimes-SR](http://hdl.handle.net/11356/1200) corpus and additional news documents from the Serbian web. @@ -13312,11 +12148,10 @@

    Serbian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13326,35 +12161,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Shanghainese + 1 + 8K + + + Sino-Tibetan, Chinese + +
    - -
    - - -

    Shanghainese treebanks

    -
    - - - -
    + +
    **UD Shanghainese-ShUD** is the first UD treebank for Shanghainese. @@ -13376,11 +12205,10 @@

    Shanghainese treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13390,35 +12218,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Sicilian + 1 + 11K + + + IE, Romance + +
    -

    Sicilian treebanks

    -
    - - - -
    + +
    The Sicilian Treebank is a small parallel corpus of Sicilian texts, automatically parsed and then manually revised, with Italian translations. It includes both contemporary and folkloric materials. The main focus is documenting typical morphosyntactic features of the written Sicilian. @@ -13440,11 +12262,10 @@

    Sicilian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13454,35 +12275,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Sindhi + 1 + 95K + + + IE, Indic + +
    - -
    - - -

    Sindhi treebanks

    -
    - - - -
    + +
    A UD dataset for Sindhi, based on newswire (primarily Kawish), folk stories from the Adabi forums, handwritten text to demonstrate @@ -13507,11 +12322,10 @@

    Sindhi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13521,35 +12335,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Sinhala + 2 + 1K + + + IE, Indic + +
    -

    Sinhala treebanks

    -
    - - - -
    + +
    This treebank contains a manually annotated Sinhala narrative based on the folk story "Appuwa", created as part of a project on treebank development for understudied languages within the Universal Dependencies framework. @@ -13571,12 +12379,11 @@

    Sinhala treebanks

  • Download
  • -

     

    - - - -
    + +
    + STB <1K @@ -13585,8 +12392,8 @@

    Sinhala treebanks

    -
    -
    + +
    This treebank consists contemporary written Sinhala text taken from a 10M corpus maintained by UCSC, Sri Lanka. The corpus contains novels, short stories, Sinhala translations, critiques and Sinhala newspapers. @@ -13598,11 +12405,10 @@

    Sinhala treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Sinhala treebanks. @@ -13614,35 +12420,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Skolt Sami + 1 + 3K + + + Uralic, Sami + +
    - -
    - - -

    Skolt Sami treebanks

    -
    - - - -
    + +
    The UD Skolt Sami Giellagas treebank is based almost entirely on spoken Skolt Sami corpora. @@ -13664,11 +12464,10 @@

    Skolt Sami treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13678,35 +12477,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Slovak + 1 + 106K + + + IE, Slavic + +
    -

    Slovak treebanks

    -
    - - - -
    + +
    The Slovak UD treebank is based on data originally annotated as part of the Slovak National Corpus, following the annotation style of the Prague @@ -13730,11 +12523,10 @@

    Slovak treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13744,35 +12536,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Slovenian + 2 + 365K + + + IE, Slavic + +
    - -
    - - -

    Slovenian treebanks

    -
    - - - -
    + +
    The SSJ treebank is the reference UD treebank for Slovenian, consisting of approximately 13,000 sentences and 267,097 tokens from fiction, non-fiction, periodical and Wikipedia texts in standard modern Slovenian. As of UD release 2.10 in May 2022, the original version of the SSJ UD treebank has been partially manually revised and extended with new manually annotated data. @@ -13794,12 +12580,11 @@

    Slovenian treebanks

  • Download
  • -

     

    - - - -
    + +
    + SST 98K @@ -13808,8 +12593,8 @@

    Slovenian treebanks

    - -
    +
    +
    The Spoken Slovenian Treebank (SST) is a manually annotated collection of transcribed audio recordings featuring spontaneous speech in various everyday situations. It includes 344 unique speech events (documents) amounting to approximately 10 hours of speech, encompassing a total of 6,121 utterances and 98,393 tokens. @@ -13821,11 +12606,10 @@

    Slovenian treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Slovenian treebanks. @@ -13837,35 +12621,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + South Levantine Arabic + 1 + <1K + + + Afro-Asiatic, Semitic + +
    -

    South Levantine Arabic treebanks

    -
    - - - -
    + +
    The South_Levantine_Arabic-MADAR treebank consists of 100 manually-annotated sentences taken from the [MADAR](https://camel.abudhabi.nyu.edu/madar/) (Multi-Arabic Dialect Applications and Resources) project. @@ -13889,11 +12667,10 @@

    South Levantine Arabic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13903,35 +12680,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Southern Kurdish + 1 + 1K + + + IE, Iranian + +
    - -
    - - -

    Southern Kurdish treebanks

    -
    - - - -
    + +
    A dependency treebank in Universal Dependencies (UD) format, derived from narrative and questionnaire texts. This document gives an overview of the annotation scheme, linguistic features, and structural patterns seen in the data. @@ -13953,11 +12724,10 @@

    Southern Kurdish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -13967,35 +12737,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Spanish + 4 + 1,022K + + + IE, Romance + +
    -

    Spanish treebanks

    -
    - - - -
    + +
    Spanish data from the [AnCora](http://clic.ub.edu/corpus/) corpus. @@ -14017,12 +12781,11 @@

    Spanish treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 23K @@ -14031,8 +12794,8 @@

    Spanish treebanks

    -
    -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -14046,12 +12809,11 @@

    Spanish treebanks

  • Download
  • -

     

    - - - - -
    + +
    The Spanish UD is converted from the content head version of the [universal dependency treebank v2.0 (legacy)](https://github.com/ryanmcd/uni-dep-tb). @@ -14074,12 +12836,11 @@

    Spanish treebanks

  • Download
  • -

     

    - - - - -
    + +
    The COSER UD Treebank (COSER-UD) is the first syntactically annotated corpus of spoken Spanish, based on a sample of the "Corpus Oral y Sonoro del Español Rural" (COSER; Fernández-Ordóñez 2005-present), meaning the "Audible Corpus of Spoken Rural Spanish". @@ -14101,11 +12862,10 @@

    Spanish treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Spanish treebanks. @@ -14117,35 +12877,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Spanish Sign Language + 1 + 1K + + + Sign Language + +
    - -
    - - -

    Spanish Sign Language treebanks

    -
    - - - -
    + +
    The Universal Dependency treebank for Spanish Sign Language (Lengua de Signos Española [LSE], ISO 639-3: ssp) was developed by the GRADES group at the University of Vigo. @@ -14167,11 +12921,10 @@

    Spanish Sign Language treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -14181,35 +12934,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Swedish + 5 + 229K + + + IE, Germanic + +
    -

    Swedish treebanks

    -
    - - - -
    + +
    UD Swedish_LinES is the Swedish half of the LinES Parallel Treebank with UD annotations. All segments are translations from English and the sources cover literary genres, online @@ -14233,12 +12980,11 @@

    Swedish treebanks

  • Download
  • -

     

    - - - -
    + +
    + Talbanken 96K @@ -14247,8 +12993,8 @@

    Swedish treebanks

    -
    -
    + +
    The Swedish-Talbanken treebank is based on Talbanken, a treebank developed at Lund University in the 1970s. @@ -14261,12 +13007,11 @@

    Swedish treebanks

  • Download
  • -

     

    - - - - -
    + +
    A treebank of learner Swedish based on SweLL, the Swedish Learner Language corpus. @@ -14288,12 +13033,11 @@

    Swedish treebanks

  • Download
  • -

     

    - - - - -
    + +
    Swedish-PUD is the Swedish part of the Parallel Universal Dependencies (PUD) treebanks. @@ -14315,12 +13059,11 @@

    Swedish treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD Swedish-Old is a treebank containing texts from Old Swedish (1225-1526). @@ -14342,11 +13085,10 @@

    Swedish treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Swedish treebanks. @@ -14358,35 +13100,29 @@

    Language documentation

    See the language documentation page. -
    - - - +
    + + - - +
    + + + Swedish Sign Language + 1 + 1K + + + Sign Language + +
    - -
    - - -

    Swedish Sign Language treebanks

    -
    - - - -
    + +
    The Universal Dependencies treebank for Swedish Sign Language (ISO 639-3: swl) is derived from the Swedish Sign Language Corpus (SSLC) from the department of @@ -14410,11 +13146,10 @@

    Swedish Sign Language treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -14424,35 +13159,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Tagalog + 2 + 1K + + + Austronesian, Greater Central Philippine + +
    -

    Tagalog treebanks

    -
    - - - -
    + +
    UD_Tagalog-TRG is a UD treebank manually annotated using sentences from a grammar book. @@ -14474,12 +13203,11 @@

    Tagalog treebanks

  • Download
  • -

     

    - - - -
    + +
    + Ugnayan 1K @@ -14488,8 +13216,8 @@

    Tagalog treebanks

    -
    -
    + +
    Ugnayan is a manually annotated Tagalog treebank currently composed of educational fiction and nonfiction text. The treebank is under development at the University of the Philippines. @@ -14501,11 +13229,10 @@

    Tagalog treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Tagalog treebanks. @@ -14517,35 +13244,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Tamil + 2 + 12K + + + Dravidian + +
    - -
    - - -

    Tamil treebanks

    -
    - - - -
    + +
    The UD Tamil treebank is based on the Tamil Dependency Treebank created at the Charles University in Prague by Loganathan Ramasamy. @@ -14568,12 +13289,11 @@

    Tamil treebanks

  • Download
  • -

     

    - - - -
    + +
    + MWTT 2K @@ -14582,8 +13302,8 @@

    Tamil treebanks

    - -
    +
    +
    MWTT - Modern Written Tamil Treebank has sentences taken primarily from a text called "A Grammar of Modern Tamil by Thomas Lehmann (1993). This initial release has 536 sentences of various lengths, and all of these are added as the test set. @@ -14595,11 +13315,10 @@

    Tamil treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Tamil treebanks. @@ -14611,35 +13330,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Tatar + 1 + 2K + + + Turkic, Northwestern + +
    -

    Tatar treebanks

    -
    - - - -
    + +
    UD Tatar-NMCTT is a manually annotated corpus of the Tatar language based on the text from Tatar-Inform (tatar-inform.tatar), an online news website. @@ -14662,11 +13375,10 @@

    Tatar treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -14676,35 +13388,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Teko + 1 + 2K + + + Tupian, Maweti-Guarani + +
    - -
    - - -

    Teko treebanks

    -
    - - - -
    + +
    UD_Teko-TuDeT is a collection of annotated sentences in <a href="https://glottolog.org/resource/languoid/id/emer1243"> Tekó (Emérillon) </a>. The sentences stem from the only grammatical description of the language (Rose, 2011). Sentence annotation and documantation by Uliana Vedenina and Fabrício Ferraz Gerardi. @@ -14726,11 +13432,10 @@

    Teko treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -14740,35 +13445,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Telugu + 1 + 6K + + + Dravidian + +
    -

    Telugu treebanks

    -
    - - - -
    + +
    The Telugu UD treebank is created in UD based on manual annotations of sentences from a grammar book. @@ -14790,11 +13489,10 @@

    Telugu treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -14804,35 +13502,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Telugu English + 1 + <1K + + + Code switching + +
    - -
    - - -

    Telugu English treebanks

    -
    - - - -
    + +
    UD Telugu_English-TECT is a Telugu-English code-switching treebank. @@ -14854,11 +13546,10 @@

    Telugu English treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -14868,35 +13559,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Thai + 2 + 99K + + + Tai-Kadai + +
    -

    Thai treebanks

    -
    - - - -
    + +
    *UD Thai TUD* (Thai Universal Dependency Treebank) is a treebank of 3,627 syntactic trees from the Thai National Corpus and Wikipedia, annotated in Universal Dependencies, covering diverse text types and topics across various domains. @@ -14918,12 +13603,11 @@

    Thai treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 22K @@ -14932,8 +13616,8 @@

    Thai treebanks

    -
    -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -14947,11 +13631,10 @@

    Thai treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Thai treebanks. @@ -14963,35 +13646,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Tswana + 1 + <1K + + + Niger-Congo, Bantoid + +
    - -
    - - -

    Tswana treebanks

    -
    - - - -
    + +
    UD Tswana-Popapolelo is a translation of the 20 Cairo Cicling sentences (https://github.com/UniversalDependencies/cairo) annotated with XPOS, UPOS and dependency relations. @@ -15013,11 +13690,10 @@

    Tswana treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15027,35 +13703,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Tupinamba + 1 + 4K + + + Tupian, Maweti-Guarani + +
    -

    Tupinamba treebanks

    -
    - - - -
    + +
    UD_Tupinamba-TuDeT is a collection of annotated sentences in [Tupinambá](https://glottolog.org/resource/languoid/id/tupi1273). All known sources in this language are being annotated: cathecisms, letters, poems, theater plays, and grammars (sixteenth and seventeenth century). Sentence annotation and documentation by [Fabrício Ferraz Gerardi](https://languagestructure.github.io). @@ -15077,11 +13747,10 @@

    Tupinamba treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15091,35 +13760,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Turkish + 10 + 735K + + + Turkic, Southwestern + +
    - -
    - - -

    Turkish treebanks

    -
    - - - -
    + +
    A Turkish dependency treebank annotated in UD style. Created by the members of [TABILAB](https://tabilab.cmpe.boun.edu.tr/) from Boğaziçi University. @@ -15141,12 +13804,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + Kenet 178K @@ -15155,8 +13817,8 @@

    Turkish treebanks

    - -
    +
    +
    Turkish-Kenet UD Treebank is the biggest treebank of Turkish. It consists of 18,700 manually annotated sentences and 178,700 tokens. Its corpus consists of dictionary examples. @@ -15168,12 +13830,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + Penn 183K @@ -15182,8 +13843,8 @@

    Turkish treebanks

    - -
    +
    +
    Turkish version of the Penn Treebank. It consists of a total of 9,560 manually annotated sentences and 87,367 tokens. (It only includes sentences up to 15 words long.) @@ -15195,12 +13856,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + Tourism 91K @@ -15209,8 +13869,8 @@

    Turkish treebanks

    - -
    +
    +
    Turkish Tourism is a domain specific treebank consisting of 19,750 manually annotated sentences and 92,200 tokens. These sentences were taken from the original customer reviews of a tourism company. @@ -15222,12 +13882,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + Atis 44K @@ -15236,8 +13895,8 @@

    Turkish treebanks

    - -
    +
    +
    This treebank is a translation of English ATIS (Airline Travel Information System) corpus (see References). It consists of 5432 sentences. @@ -15249,12 +13908,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + IMST 58K @@ -15263,8 +13921,8 @@

    Turkish treebanks

    - -
    +
    +
    The UD Turkish Treebank, also called the IMST-UD Treebank, is a semi-automatic conversion of the IMST Treebank (Sulubacak&Eryiğit, 2018; Sulubacak et al., 2016). @@ -15276,12 +13934,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + FrameNet 19K @@ -15290,8 +13947,8 @@

    Turkish treebanks

    - -
    +
    +
    Turkish FrameNet consists of 2,700 manually annotated example sentences and 19,221 tokens. Its data consists of the sentences taken from the Turkish FrameNet Project. The annotated sentences can be filtered according to the semantic frame category of the root of the sentence. @@ -15303,12 +13960,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL <1K @@ -15317,8 +13973,8 @@

    Turkish treebanks

    - -
    +
    +
    The Turkish-TueCL treebank is part of a parallel Universal Dependencies corpus containing 148 sentences across four Turkic languages (Turkish, Azerbaijani, Kyrgyz, and Uzbek), designed to facilitate cross-linguistic research on these related languages. @@ -15330,12 +13986,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + GB 17K @@ -15344,8 +13999,8 @@

    Turkish treebanks

    - -
    +
    +
    This is a treebank annotating example sentences from a comprehensive grammar book of Turkish. @@ -15357,12 +14012,11 @@

    Turkish treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD 16K @@ -15371,8 +14025,8 @@

    Turkish treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -15386,11 +14040,10 @@

    Turkish treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Turkish treebanks. @@ -15402,35 +14055,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Turkish English + 1 + <1K + + + Code switching + +
    -

    Turkish English treebanks

    -
    - - - -
    + +
    UD_Turkish_English-BUTR is a treebank of Turkish-English code-switched sentences collected from Boğaziçi University students, annotated in the Universal Dependencies framework to provide a standardized resource for analyzing syntactic patterns in Turkish-English code-switching. @@ -15452,11 +14099,10 @@

    Turkish English treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15466,35 +14112,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Turkish German + 1 + 37K + + + Code switching + +
    - -
    - - -

    Turkish German treebanks

    -
    - - - -
    + +
    UD Turkish-German SAGT is a Turkish-German code-switching treebank that is developed as part of the [SAGT](https://www.ims.uni-stuttgart.de/en/research/projects/sagt/) project. @@ -15516,11 +14156,10 @@

    Turkish German treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15530,35 +14169,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Ukrainian + 2 + 231K + + + IE, Slavic + +
    -

    Ukrainian treebanks

    -
    - - - -
    + +
    UD_Ukrainian-ParlaMint is a collection of Ukrainian parliamentary transcripts annotated in Universal Dependencies. The texts are published on the official website of the Ukrainian parliament (https://www.rada.gov.ua/documents/Stenbul_pz/) and are taken for UD_Ukrainian-ParlaMint from the Ukrainian section of the ParlaMint project (https://www.clarin.eu/parlamint). @@ -15580,12 +14213,11 @@

    Ukrainian treebanks

  • Download
  • -

     

    - - - -
    + +
    + IU 122K @@ -15594,8 +14226,8 @@

    Ukrainian treebanks

    -
    -
    + +
    Gold standard Universal Dependencies corpus for Ukrainian, developed for UD originally, by [Institute for Ukrainian](https://mova.institute), NGO. [[українською](https://mova.institute/золотий_стандарт)] @@ -15607,11 +14239,10 @@

    Ukrainian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Ukrainian treebanks. @@ -15623,35 +14254,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Umbrian + 1 + 1K + + + IE, Italic + +
    - -
    - - -

    Umbrian treebanks

    -
    - - - -
    + +
    UD_Umbrian-IKUVINA is a dependency treebank rendering of the Iguvine tablets ([Wikipedia](https://en.wikipedia.org/wiki/Iguvine_Tablets)). The seven bronze tablets describe religious ceremonies performed by the Umbrian people in Italy before the rise of the Roman empire. @@ -15677,11 +14302,10 @@

    Umbrian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15691,35 +14315,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Upper Sorbian + 1 + 11K + + + IE, Slavic + +
    -

    Upper Sorbian treebanks

    -
    - - - -
    + +
    A small treebank of Upper Sorbian based mostly on Wikipedia, partly also on other Sorbian websites. @@ -15741,11 +14359,10 @@

    Upper Sorbian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15755,35 +14372,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Urdu + 1 + 138K + + + IE, Indic + +
    - -
    - - -

    Urdu treebanks

    -
    - - - -
    + +
    The Urdu Universal Dependency Treebank was automatically converted from Urdu Dependency Treebank (UDTB) which is part of an ongoing effort of creating multi-layered treebanks for Hindi and Urdu. @@ -15805,11 +14416,10 @@

    Urdu treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15819,35 +14429,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Uyghur + 1 + 40K + + + Turkic, Southeastern + +
    -

    Uyghur treebanks

    -
    - - - -
    + +
    The Uyghur UD treebank is based on the Uyghur Dependency Treebank (UDT), created at the Xinjiang University in Ürümqi, China. @@ -15870,11 +14474,10 @@

    Uyghur treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -15884,35 +14487,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Uzbek + 3 + 14K + + + Turkic, Southeastern + +
    - -
    - - -

    Uzbek treebanks

    -
    - - - -
    + +
    UD_Uzbek-UzUDT is a manually annotated Universal Dependencies treebank for Uzbek language. @@ -15934,12 +14531,11 @@

    Uzbek treebanks

  • Download
  • -

     

    - - - -
    + +
    + UT 5K @@ -15948,8 +14544,8 @@

    Uzbek treebanks

    - -
    +
    +
    This is the first Uzbek UD treebank. @@ -15961,12 +14557,11 @@

    Uzbek treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL <1K @@ -15975,8 +14570,8 @@

    Uzbek treebanks

    - -
    +
    +
    The Uzbek-TueCL treebank is part of a parallel Universal Dependencies corpus containing 148 sentences across four Turkic languages: Turkish, Azerbaijani, Kyrgyz, and Uzbek. @@ -15988,11 +14583,10 @@

    Uzbek treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Uzbek treebanks. @@ -16004,35 +14598,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Veps + 1 + 1K + + + Uralic, Finnic + +
    -

    Veps treebanks

    -
    - - - -
    + +
    UD Veps-VWT is a manually annotated corpus of Veps made using the Universal dependencies annotation scheme. The data is collected from [VepKar corpora](http://dictorpus.krc.karelia.ru/en/corpus/text) and consists of mostly modern news texts written in Central Veps dialect. @@ -16054,11 +14642,10 @@

    Veps treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16068,35 +14655,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Vietnamese + 2 + 59K + + + Austro-Asiatic, Viet-Muong + +
    - -
    - - -

    Vietnamese treebanks

    -
    - - - -
    + +
    The Vietnamese UD treebank is a conversion of the constituent treebank created in the VLSP project (https://vlsp.hpda.vn/). @@ -16119,12 +14700,11 @@

    Vietnamese treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL 1K @@ -16133,8 +14713,8 @@

    Vietnamese treebanks

    - -
    +
    +
    This treebank includes a set of sentences from [OPUS](https://opus.nlpl.eu/), sourced from subtitles, talks, and educational videos. @@ -16147,11 +14727,10 @@

    Vietnamese treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Vietnamese treebanks. @@ -16163,35 +14742,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Warlpiri + 1 + <1K + + + Pama-Nyungan, Western + +
    -

    Warlpiri treebanks

    -
    - - - -
    + +
    A small treebank of grammatical examples in Warlpiri, taken from linguistic literature. @@ -16213,11 +14786,10 @@

    Warlpiri treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16227,35 +14799,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Welsh + 1 + 54K + + + IE, Celtic + +
    - -
    - - -

    Welsh treebanks

    -
    - - - -
    + +
    UD Welsh-CCG (Corpws Cystrawennol y Gymraeg) is a treebank of Welsh, annotated according to the Universal Dependencies guidelines. @@ -16278,11 +14844,10 @@

    Welsh treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16292,35 +14857,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Western Armenian + 1 + 122K + + + IE, Armenian + +
    -

    Western Armenian treebanks

    -
    - - - -
    + +
    A Universal Dependencies treebank for Western Armenian was developed for UD originally by the ArmTDP team led by Marat M. Yavrumyan at the Yerevan State University. @@ -16342,11 +14901,10 @@

    Western Armenian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16356,35 +14914,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Western S.P. Nahuatl + 1 + 19K + + + Uto-Aztecan + +
    - -
    - - -

    Western Sierra Puebla Nahuatl treebanks

    -
    - - - -
    + +
    UD Western Sierra Puebla Nahuatl-MesoTree is a combination of the existing UD Western Sierra Puebla Nahuatl-IU treebank (ITML) (with some updates to annotations due to caught errors or changes annotation decisions) and new sentences annotated as part of the NSF-funded project, "Syntactically-annotated corpora for endangered languages in areal contact" (MesoTree). @@ -16406,11 +14958,10 @@

    Western Sierra Puebla Nahuatl treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16420,35 +14971,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Wolof + 1 + 44K + + + Niger-Congo, Northern Atlantic + +
    -

    Wolof treebanks

    -
    - - - -
    + +
    UD_Wolof-WTB is a natively manual developed treebank for Wolof. Sentences were collected from encyclopedic, fictional, biographical, religious texts and news. @@ -16470,11 +15015,10 @@

    Wolof treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16484,35 +15028,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Xavante + 1 + 2K + + + Macro-Je + +
    - -
    - - -

    Xavante treebanks

    -
    - - - -
    + +
    UD_Xavante-XDT is a collection of annotated sentences in [Xavante](https://glottolog.org/resource/languoid/id/xava1240). Sentence annotation and documentation by [Fabrício Ferraz Gerardi](http://languagestructure.github.io/), Ivan Roksandic. @@ -16534,11 +15072,10 @@

    Xavante treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16548,35 +15085,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Xibe + 1 + 15K + + + Tungusic + +
    -

    Xibe treebanks

    -
    - - - -
    + +
    The UD Xibe Treebank is a corpus of the Xibe language (ISO 639-3: *sjo*) containing manually annotated syntactic trees under the Universal Dependencies. Sentences come from three sources: grammar book examples, newspaper (Cabcal News) and Xibe textbooks. @@ -16598,11 +15129,10 @@

    Xibe treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16612,35 +15142,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Yakut + 1 + 1K + + + Turkic, Northeastern + +
    - -
    - - -

    Yakut treebanks

    -
    - - - -
    + +
    UD_Yakut-YKTDT is a collection Yakut ([Sakha]) sentences (https://glottolog.org/resource/languoid/id/yaku1245). The project is work-in-progress and the treebank is being updated on a regular basis. @@ -16662,11 +15186,10 @@

    Yakut treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16676,35 +15199,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Yiddish + 1 + 28K + + + IE, Germanic + +
    -

    Yiddish treebanks

    -
    - - - -
    + +
    YiTB is a treebank of linguistically annotated Yiddish data in the Universal Dependencies framework, created via a bootstraping machine learning method. A total of 27,872 tokens are currently in the treebank from a variety of sources and textual genres. @@ -16726,11 +15243,10 @@

    Yiddish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16740,35 +15256,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Yoruba + 1 + 8K + + + Niger-Congo, Defoid + +
    - -
    - - -

    Yoruba treebanks

    -
    - - - -
    + +
    Parts of the Yoruba Bible and of the Yoruba edition of Wikipedia, hand-annotated natively in Universal Dependencies. @@ -16790,11 +15300,10 @@

    Yoruba treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16804,35 +15313,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Yupik + 1 + 2K + + + Eskimo-Aleut + +
    -

    Yupik treebanks

    -
    - - - -
    + +
    UD_Yupik-SLI is a treebank of St. Lawrence Island Yupik (ISO 639-3: ess) that has been manually annotated at the morpheme level, based on a finite-state morphological analyzer by [Chen et al., 2020](https://www.aclweb.org/anthology/2020.lrec-1.326). The word-level annotation, merging multiword expressions, is provided in not-to-release/ess_sli-ud-test.merged.conllu. @@ -16856,11 +15359,10 @@

    Yupik treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16870,35 +15372,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Zaar + 1 + 20K + + + Afro-Asiatic, West Chadic + +
    - -
    - - -

    Zaar treebanks

    -
    - - - -
    + +
    A Universal Dependencies corpus for Zaar (aka Sayanci), a member of the Chadic branch of the Afro-Asiatic phylum. The language is mainly spoken by about 200,000 speakers in the Bogoro and Tafawa Balewa local governments of Bauchi State, Nigeria. @@ -16920,11 +15416,10 @@

    Zaar treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16934,35 +15429,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Zazaki + 1 + 1K + + + IE, Iranian + +
    -

    Zazaki treebanks

    -
    - - - -
    + +
    A dependency treebank in Universal Dependencies (UD) format for Zazakî (Kirmanckî), based on transcribed spoken interviews recorded in situ in Dêrsim (Turk. Tunceli) for the forthcoming documentary "KUTENE – Last Dots of Dêrsim". It provides manually verified POS tags, morphological features, and dependency relations. @@ -16984,11 +15473,10 @@

    Zazaki treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -16998,6 +15486,7 @@

    Language documentation

    See the language documentation page. -
    - + + + diff --git a/_includes/at_glance_retired.html b/_includes/at_glance_retired.html index b85882d2f9..d445674dc1 100644 --- a/_includes/at_glance_retired.html +++ b/_includes/at_glance_retired.html @@ -1,29 +1,22 @@ - - - - - - -
    - +
    + + + English + 1 + 97K + + + IE, Germanic + +
    -

    English treebanks

    -
    - - - -
    + +
    UD English-ESL / Treebank of Learner English (TLE) contains manual POS tag and dependency annotations for 5,124 English as a Second Language (ESL) sentences drawn from the Cambridge Learner Corpus First Certificate in English (FCE) dataset. @@ -45,11 +38,10 @@

    English treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of English treebanks. @@ -61,35 +53,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + French + 1 + 573K + + + IE, Romance + +
    - -
    - - -

    French treebanks

    -
    - - - -
    + +
    The Universal Dependency version of the French Treebank (Abeillé et al., 2003), hereafter UD_French-FTB, is a treebank of sentences from the newspaper Le Monde, initially manually annotated with morphological information and phrase-structure and then converted to the Universal Dependencies annotation scheme. @@ -111,11 +97,10 @@

    French treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of French treebanks. @@ -127,35 +112,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Hindi English + 1 + 26K + + + Code switching + +
    -

    Hindi English treebanks

    -
    - - - -
    + +
    The Hindi-English Code-switching treebank is based on code-switching tweets of Hindi and English multilingual speakers (mostly Indian) on Twitter. The treebank is manually annotated using UD sceheme. The training and evaluations sets were seperately annotated by different annotators using UD v2 and v1 guidelines respectively. The evaluation sets are automatically converted from UD v1 to v2. @@ -177,11 +156,10 @@

    Hindi English treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -191,35 +169,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Japanese + 2 + 204K + + + Japanese + +
    - -
    - - -

    Japanese treebanks

    -
    - - - -
    + +
    This Universal Dependencies (UD) Japanese treebank is based on the definition of UD Japanese convention described in the UD documentation. @@ -243,12 +215,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - -
    + +
    + KTC 189K @@ -257,8 +228,8 @@

    Japanese treebanks

    - -
    +
    +
    Please add a summary section to the treebank readme file @@ -270,11 +241,10 @@

    Japanese treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Japanese treebanks. @@ -286,35 +256,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Khunsari + 1 + <1K + + + IE, Iranian + +
    -

    Khunsari treebanks

    -
    - - - -
    + +
    The AHA Khunsari Treebank is a small treebank for contemporary Khunsari. Its corpus is collected and annotated manually. We have prepared this treebank based on interviews with Khunsari speakers. @@ -336,11 +300,10 @@

    Khunsari treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -350,35 +313,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Mbya Guarani + 1 + 11K + + + Tupian, Maweti-Guarani + +
    - -
    - - -

    Mbya Guarani treebanks

    -
    - - - -
    + +
    UD Mbya_Guarani-Dooley is a corpus of narratives written in Mbyá Guaraní (Tupian) in Brazil, and collected by Robert Dooley. Due to copyright restrictions, the corpus that is distributed as part of UD only contains the annotation (tags, features, relations) while the FORM and LEMMA columns are empty. @@ -400,11 +357,10 @@

    Mbya Guarani treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Mbya Guarani treebanks. @@ -416,35 +372,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Nayini + 1 + <1K + + + IE, Iranian + +
    -

    Nayini treebanks

    -
    - - - -
    + +
    The AHA Nayini Treebank is a small treebank for contemporary Nayini. Its corpus is collected and annotated manually. We have prepared this treebank based on interviews with Nayini speakers. @@ -466,11 +416,10 @@

    Nayini treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -480,35 +429,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Norwegian + 1 + 55K + + + IE, Germanic + +
    - -
    - - -

    Norwegian treebanks

    -
    - - - -
    + +
    This Norwegian treebank is based on the LIA treebank of transcribed spoken Norwegian dialects. The treebank has been automatically converted to the UD scheme by Lilja Øvrelid at the University of Oslo. @@ -531,11 +474,10 @@

    Norwegian treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Norwegian treebanks. @@ -547,35 +489,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Soi + 1 + <1K + + + IE, Iranian + +
    -

    Soi treebanks

    -
    - - - -
    + +
    The AHA Soi Treebank is a small treebank for contemporary Soi. Its corpus is collected and annotated manually. We have prepared this treebank based on interviews with Soi speakers. @@ -597,11 +533,10 @@

    Soi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -611,6 +546,7 @@

    Language documentation

    See the language documentation page. -
    - + + + diff --git a/_includes/at_glance_sapling.html b/_includes/at_glance_sapling.html index 60db6ae7c8..8d083cbe39 100644 --- a/_includes/at_glance_sapling.html +++ b/_includes/at_glance_sapling.html @@ -1,29 +1,22 @@ - - - - +
    + + + Akkadian + 1 + 117K + + + Afro-Asiatic, Semitic + +
    - -
    - - -

    Akkadian treebanks

    -
    - - - -
    + +
    UD_Akkadian-MCONG is a treebank of normalized Akkadian sentences drawn mostly from Neo-Assyrian corpora lemmatized on [Oracc](http://oracc.museum.upenn.edu/). Sentences are annotated for lemma, syntactic dependencies, and morphological features. The treebank contains approximately 112,000 words. @@ -45,11 +38,10 @@

    Akkadian treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Akkadian treebanks. @@ -61,35 +53,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Amharic + 3 + <1K + + + Afro-Asiatic, Semitic + +
    -

    Amharic treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -111,12 +97,11 @@

    Amharic treebanks

  • Download
  • -

     

    - - - -
    + +
    + Inku - @@ -125,8 +110,8 @@

    Amharic treebanks

    -
    -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -138,12 +123,11 @@

    Amharic treebanks

  • Download
  • -

     

    - - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -165,11 +149,10 @@

    Amharic treebanks

  • Download
  • -

     

    - - -
    + + + @@ -179,35 +162,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Apalai + 1 + <1K + + + Cariban + +
    - -
    - - -

    Apalai treebanks

    -
    - - - -
    + +
    This Apalaí treebank is a collection of sentences from different sources, including own fieldwork material. @@ -229,11 +206,10 @@

    Apalai treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -243,35 +219,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Balatipone + 1 + - + + + Bororoan + +
    -

    Balatipone treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -293,11 +263,10 @@

    Balatipone treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -307,35 +276,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Balochi + 1 + - + + + IE, Iranian + +
    - -
    - - -

    Balochi treebanks

    -
    - - - -
    + +
    UD_Balochi-GPS is a treebank of the Balochi language variety spoken in southeastern Iran annotated according to the Universal Dependencies framework. @@ -357,11 +320,10 @@

    Balochi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -371,35 +333,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Bengali + 3 + 2K + + + IE, Indic + +
    -

    Bengali treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -421,12 +377,11 @@

    Bengali treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD - @@ -435,8 +390,8 @@

    Bengali treebanks

    -
    -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks originally created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -450,12 +405,11 @@

    Bengali treebanks

  • Download
  • -

     

    - - - - -
    + +
    UD_Bengali-Sabdakosh is a corpus parsed sentences from contemporary Bengali prose, consisting of passages of up to 50 sentences from modern fiction. @@ -477,11 +431,10 @@

    Bengali treebanks

  • Download
  • -

     

    - - -
    + + + @@ -491,35 +444,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Central Romani + 1 + - + + + IE, Indic + +
    - -
    - - -

    Central Romani treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -541,11 +488,10 @@

    Central Romani treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -555,35 +501,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Classical Nahuatl + 1 + - + + + Uto-Aztecan + +
    -

    Classical Nahuatl treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -605,11 +545,10 @@

    Classical Nahuatl treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -619,35 +558,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Corsican + 1 + - + + + IE, Romance + +
    - -
    - - -

    Corsican treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -669,11 +602,10 @@

    Corsican treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -683,35 +615,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Cuicatec + 1 + - + + + Oto-Manguean + +
    -

    Cuicatec treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -733,11 +659,10 @@

    Cuicatec treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -747,35 +672,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Cusco Quechua + 1 + - + + + Quechuan + +
    - -
    - - -

    Cusco Quechua treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -797,11 +716,10 @@

    Cusco Quechua treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -811,35 +729,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Czech + 1 + - + + + IE, Slavic + +
    -

    Czech treebanks

    -
    - - - -
    + +
    Morphosyntactic annotation of 1600 sentences from the Czesl-MAN corpus (Czech as a second language). @@ -861,11 +773,10 @@

    Czech treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Czech treebanks. @@ -877,35 +788,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Dargwa + 1 + - + + + Nakh-Daghestanian, Lak-Dargwa + +
    - -
    - - -

    Dargwa treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -927,11 +832,10 @@

    Dargwa treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -941,35 +845,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Enawene Nawe + 1 + - + + + Arawakan, Central Arawakan + +
    -

    Enawene Nawe treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -991,11 +889,10 @@

    Enawene Nawe treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1005,35 +902,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + French + 1 + - + + + IE, Romance + +
    - -
    - - -

    French treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1055,11 +946,10 @@

    French treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of French treebanks. @@ -1071,35 +961,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + French Sign Language + 1 + - + + + Sign Language + +
    -

    French Sign Language treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1121,11 +1005,10 @@

    French Sign Language treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1135,35 +1018,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Frisian + 1 + 51K + + + IE, Germanic + +
    - -
    - - -

    Frisian treebanks

    -
    - - - -
    + +
    The UD Frisian-FA-RuG treebank is a West Frisian treebank. @@ -1185,11 +1062,10 @@

    Frisian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1199,35 +1075,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Gedeo + 1 + <1K + + + Afro-Asiatic, Cushitic + +
    -

    Gedeo treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1249,11 +1119,10 @@

    Gedeo treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1263,35 +1132,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Georgian + 1 + 1K + + + Kartvelian + +
    - -
    - - -

    Georgian treebanks

    -
    - - - -
    + +
    **UD\_Georgian-GEOWIKI** is a Universal Dependencies corpus for the Georgian language, derived from randomly selected texts from the Georgian Wikipedia. The corpus currently contains 385 sentences, which were automatically tokenized and morphologically tagged using a TreeTagger model trained on a separate Georgian dataset. You can find the TreeTagger resource [here](https://github.com/SophikoComp/TreeTagger-for-Georgian). @@ -1317,11 +1180,10 @@

    Georgian treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Georgian treebanks. @@ -1333,35 +1195,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Gilaki + 1 + - + + + IE, Iranian + +
    -

    Gilaki treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1383,11 +1239,10 @@

    Gilaki treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1397,35 +1252,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Greek + 2 + - + + + IE, Greek + +
    - -
    - - -

    Greek treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -1447,12 +1296,11 @@

    Greek treebanks

  • Download
  • -

     

    - - - -
    + +
    + Griko - @@ -1461,8 +1309,8 @@

    Greek treebanks

    - -
    +
    +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1474,11 +1322,10 @@

    Greek treebanks

  • Download
  • -

     

    - - - +
    + + See here for comparative statistics of Greek treebanks. @@ -1490,35 +1337,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Hiligaynon + 1 + <1K + + + Austronesian, Greater Central Philippine + +
    -

    Hiligaynon treebanks

    -
    - - - -
    + +
    UD Hiligaynon-HTB is a UD treebank containing sentences manually-annotated from grammar books [PALI Language Texts](https://www.hawaiiopen.org/bookseries/pali-language-texts-philippines/) @@ -1542,11 +1383,10 @@

    Hiligaynon treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1556,35 +1396,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Hindi + 1 + 4K + + + IE, Indic + +
    - -
    - - -

    Hindi treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1606,11 +1440,10 @@

    Hindi treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Hindi treebanks. @@ -1622,35 +1455,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Huave + 1 + - + + + Huavean + +
    -

    Huave treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1672,11 +1499,10 @@

    Huave treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1686,35 +1512,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Italian + 1 + - + + + IE, Romance + +
    - -
    - - -

    Italian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -1736,11 +1556,10 @@

    Italian treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Italian treebanks. @@ -1752,35 +1571,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Japanese + 2 + - + + + Japanese + +
    -

    Japanese treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1802,12 +1615,11 @@

    Japanese treebanks

  • Download
  • -

     

    - - - -
    + +
    + JDDLUW - @@ -1816,8 +1628,8 @@

    Japanese treebanks

    -
    -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1829,11 +1641,10 @@

    Japanese treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Japanese treebanks. @@ -1845,35 +1656,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Kabyle + 1 + 23K + + + Afro-Asiatic, Berber + +
    - -
    - - -

    Kabyle treebanks

    -
    - - - -
    + +
    UD UD_Kabyle-ADPT (Association pour le Développement et la Promotion de Tamazight) is a treebank of Berber (Kabyle variant), annotated according to the Universal Dependencies guidelines. @@ -1896,11 +1701,10 @@

    Kabyle treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1910,35 +1714,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Kiga + 1 + - + + + Niger-Congo, Bantoid + +
    -

    Kiga treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -1960,11 +1758,10 @@

    Kiga treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -1974,35 +1771,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Komi + 1 + <1K + + + Uralic, Permic + +
    - -
    - - -

    Komi treebanks

    -
    - - - -
    + +
    This is an Universal Dependencies treebank of Old Permic. The treebank is currently under progress, and will be published completely in the next Universal Dependencies release (v2.14), which is scheduled for May 15, 2024 (data freeze on May 1). @@ -2024,11 +1815,10 @@

    Komi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2038,35 +1828,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Kullvi + 1 + - + + + IE, Indic + +
    -

    Kullvi treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -2088,11 +1872,10 @@

    Kullvi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2102,35 +1885,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Ladino + 1 + - + + + IE, Romance + +
    - -
    - - -

    Ladino treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2152,11 +1929,10 @@

    Ladino treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2166,35 +1942,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Laz + 1 + 2K + + + Kartvelian + +
    -

    Laz treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2216,11 +1986,10 @@

    Laz treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2230,35 +1999,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Magahi + 2 + 7K + + + IE, Indic + +
    - -
    - - -

    Magahi treebanks

    -
    - - - -
    + +
    The [Magahi](https://en.wikipedia.org/wiki/Magahi_language) UD Treebank (MGTB) is a part of the [Universal Dependency treebank](http://universaldependencies.org/) project. @@ -2280,12 +2043,11 @@

    Magahi treebanks

  • Download
  • -

     

    - - - -
    + +
    + PUD - @@ -2294,8 +2056,8 @@

    Magahi treebanks

    - -
    +
    +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks originally created for the [CoNLL 2017 shared task on Multilingual Parsing from Raw Text to @@ -2309,11 +2071,10 @@

    Magahi treebanks

  • Download
  • -

     

    - - - +
    + + @@ -2323,35 +2084,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Malagasy + 1 + - + + + Austronesian, Barito + +
    -

    Malagasy treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2373,11 +2128,10 @@

    Malagasy treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2387,35 +2141,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Mandyali + 1 + 2K + + + IE, Indic + +
    - -
    - - -

    Mandyali treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2437,11 +2185,10 @@

    Mandyali treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2451,35 +2198,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Mansi + 1 + - + + + Uralic, Ugric + +
    -

    Mansi treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2501,11 +2242,10 @@

    Mansi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2515,35 +2255,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Megrelian + 1 + - + + + Kartvelian + +
    -

    Megrelian treebanks

    -
    - - - -
    + +
    The Megrelian UD Treebank (UD_Megrelian-MLC) is the first syntactically annotated corpus of Megrelian, based on a collection of annotated sentences selected from the Megrelian Language Corpus (MLC) available at http://xmf.iliauni.edu.ge/ . @@ -2565,11 +2299,10 @@

    Megrelian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2579,35 +2312,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Middle Irish + 2 + <1K + + + IE, Celtic + +
    - -
    - - -

    Middle Irish treebanks

    -
    - - - -
    + +
    Annotation of the classic Scela Mucce Meic Dathó ("The tale of Mac Dathó's pig"). @@ -2629,12 +2356,11 @@

    Middle Irish treebanks

  • Download
  • -

     

    - - - -
    + +
    + DipMITB - @@ -2643,8 +2369,8 @@

    Middle Irish treebanks

    - -
    +
    +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2656,11 +2382,10 @@

    Middle Irish treebanks

  • Download
  • -

     

    - - - +
    + + @@ -2670,35 +2395,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Middle Persian + 1 + - + + + IE, Iranian + +
    -

    Middle Persian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2720,11 +2439,10 @@

    Middle Persian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2734,35 +2452,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Mongolian + 1 + - + + + Mongolic + +
    - -
    - - -

    Mongolian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2784,11 +2496,10 @@

    Mongolian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2798,35 +2509,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - +
    + + + Nkore + 1 + - + + + Niger-Congo, Bantoid + +
    - -
    - - -

    Nkore treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -2848,11 +2553,10 @@

    Nkore treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2862,35 +2566,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Northern Kurdish + 1 + 2K + + + IE, Iranian + +
    -

    Northern Kurdish treebanks

    -
    - - - -
    + +
    The Bezeyni treebank is a small collection of annotated sentences in Bezeynî, a Kurdish language variety spoken in Turkey. The corpus contains 177 sentences with morphological and syntactic annotations following Universal Dependencies guidelines, providing basic coverage of nominal, verbal, and clause structures. @@ -2912,11 +2610,10 @@

    Northern Kurdish treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2926,35 +2623,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Northern Luri + 1 + - + + + IE, Iranian + +
    - -
    - - -

    Northern Luri treebanks

    -
    - - - -
    + +
    UD_NorthernLuri-KHS is a treebank of the Khoranabadi variety of Northern Luri, annotated according to the Universal Dependencies framework. @@ -2976,11 +2667,10 @@

    Northern Luri treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -2990,35 +2680,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Occitan + 1 + - + + + IE, Romance + +
    - -
    - - -

    Occitan treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3040,11 +2724,10 @@

    Occitan treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3054,35 +2737,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old English + 2 + 26K + + + IE, Germanic + +
    -

    Old English treebanks

    -
    - - - -
    + +
    This is a 25,000 word UD treebank of Old English. The text has been retrieved from Martín Arista, Javier (ed.), et al. 2023. ParCorOEv3 [www.nerthusproject.com]. @@ -3106,12 +2783,11 @@

    Old English treebanks

  • Download
  • -

     

    - - - -
    + +
    + TueCL - @@ -3120,8 +2796,8 @@

    Old English treebanks

    -
    -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3133,11 +2809,10 @@

    Old English treebanks

  • Download
  • -

     

    - - -
    + + + @@ -3147,35 +2822,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Old Georgian + 2 + - + + + Kartvelian + +
    - -
    - - -

    Old Georgian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -3197,12 +2866,11 @@

    Old Georgian treebanks

  • Download
  • -

     

    - - - -
    + +
    + OGLauRo - @@ -3211,8 +2879,8 @@

    Old Georgian treebanks

    - -
    +
    +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -3224,11 +2892,10 @@

    Old Georgian treebanks

  • Download
  • -

     

    - - - +
    + + @@ -3238,35 +2905,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old Japanese + 1 + 3K + + + Japanese + +
    -

    Old Japanese treebanks

    -
    - - - -
    + +
    UD_Old_Japanese-LMJ is a collection of annotated texts in Late Middle Japanese, starting with Book 9 from he celebrated gunki monogatari (war tale) *Heike Monogatari*. @@ -3288,11 +2949,10 @@

    Old Japanese treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3302,35 +2962,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Old Occitan + 1 + - + + + IE, Romance + +
    - -
    - - -

    Old Occitan treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3352,11 +3006,10 @@

    Old Occitan treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3366,35 +3019,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Old Saxon + 1 + - + + + IE, Germanic + +
    -

    Old Saxon treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3416,11 +3063,10 @@

    Old Saxon treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3430,35 +3076,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Palenquero + 1 + - + + + Creole + +
    -

    Palenquero treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3480,11 +3120,10 @@

    Palenquero treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3494,35 +3133,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Pali + 1 + - + + + IE, Indic + +
    - -
    - - -

    Pali treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -3544,11 +3177,10 @@

    Pali treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3558,35 +3190,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Papiamento + 2 + - + + + Creole + +
    -

    Papiamento treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3608,12 +3234,11 @@

    Papiamento treebanks

  • Download
  • -

     

    - - - -
    + +
    + CW - @@ -3622,8 +3247,8 @@

    Papiamento treebanks

    -
    -
    + +
    If you can read this sentence, then we are still working on our first release. @@ -3635,11 +3260,10 @@

    Papiamento treebanks

  • Download
  • -

     

    - - -
    + + + @@ -3649,35 +3273,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - +
    + + + Peripheral Mongolian + 1 + - + + + Mongolic + +
    - -
    - - -

    Peripheral Mongolian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](https://universaldependencies.org/contributing/repository_files.html#the-readme-file) for README guidelines) ... @@ -3699,11 +3317,10 @@

    Peripheral Mongolian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3713,35 +3330,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Persian + 3 + - + + + IE, Iranian + +
    -

    Persian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3763,12 +3374,11 @@

    Persian treebanks

  • Download
  • -

     

    - - - -
    + +
    + IPerUDT - @@ -3777,8 +3387,8 @@

    Persian treebanks

    -
    -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3790,12 +3400,11 @@

    Persian treebanks

  • Download
  • -

     

    - - - - -
    + +
    This is a part of the Parallel Universal Dependencies (PUD) treebanks (original set of languages annotated for CoNLL 2017 shared task; @@ -3819,11 +3428,10 @@

    Persian treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Persian treebanks. @@ -3835,35 +3443,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - +
    + + + Pnar + 1 + - + + + Austro-Asiatic, Khasian + +
    - -
    - - -

    Pnar treebanks

    -
    - - - -
    + +
    UD Pnar-PTB is a conversion from the Ring (2017) dataset ([doi:10.21979/N9/KVFGBZ](http://dx.doi.org/10.21979/N9/KVFGBZ)) that underpins a grammatical description of the Pnar language (Ring 2015, [http://hdl.handle.net/10356/62519](http://hdl.handle.net/10356/62519)). The corpus consists of folktales and interviews transcribed, translated, and interlinearized. @@ -3885,11 +3487,10 @@

    Pnar treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3899,35 +3500,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - +
    + + + Pontic + 1 + - + + + IE, Greek + +
    - -
    - - -

    Pontic treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -3949,11 +3544,10 @@

    Pontic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -3963,35 +3557,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - - - -
    - +
    + + + Portuguese + 2 + 101K + + + IE, Romance + +
    -

    Portuguese treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4013,12 +3601,11 @@

    Portuguese treebanks

  • Download
  • -

     

    - - - -
    + +
    + PortJur 101K @@ -4027,8 +3614,8 @@

    Portuguese treebanks

    -
    -
    + +
    The legal portion of [Porttinari](https://sites.google.com/icmc.usp.br/poetisa/resources-and-tools), which includes public law texts produced by the judiciary (mainly summaries) and the legislature (laws), including widely known laws in Brazil, as Henry Borel law, Internet Civil Rights law, Maria da Penha law, Copyright law, Agrarian Reform law, Elderly Persons statute and Child and Adolescent statute. @@ -4040,11 +3627,10 @@

    Portuguese treebanks

  • Download
  • -

     

    - - -
    + + + See here for comparative statistics of Portuguese treebanks. @@ -4056,35 +3642,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - +
    + + + Prakrit + 1 + <1K + + + IE, Indic + +
    - -
    - - -

    Prakrit treebanks

    -
    - - - -
    + +
    **UD Prakrit-DIPI** (*Digitising Imperial Prakrit Inscriptions*) is a UD-annotated corpus of the Ashokan Prakrit inscriptions and edicts (parallel texts written in various dialects) representing an early stage of Middle Indo-Aryan. This corpus aims to facilitate comparative work on the Ashokan dialects with the help of new computational methods. @@ -4106,11 +3686,10 @@

    Prakrit treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4120,35 +3699,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Punjabi + 1 + 6K + + + IE, Indic + +
    -

    Punjabi treebanks

    -
    - - - -
    + +
    **PunTB** (a very imaginative acronym for **Pun**jabi **T**ree**b**ank) is an in-progress treebank of Punjabi in the Gurmukhi script, aiming to cover a wide range of genres and formats. @@ -4170,11 +3743,10 @@

    Punjabi treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Punjabi treebanks. @@ -4186,35 +3758,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Puno Quechua + 1 + - + + + Quechuan + +
    -

    Puno Quechua treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4236,11 +3802,10 @@

    Puno Quechua treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4250,35 +3815,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Sardinian + 2 + - + + + IE, Romance + +
    - -
    - - -

    Sardinian treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4300,12 +3859,11 @@

    Sardinian treebanks

  • Download
  • -

     

    - - - -
    + +
    + EModSar - @@ -4314,8 +3872,8 @@

    Sardinian treebanks

    - -
    +
    +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4327,11 +3885,10 @@

    Sardinian treebanks

  • Download
  • -

     

    - - - +
    + + @@ -4341,35 +3898,29 @@

    Language documentation

    See the language documentation page. - - - - + + + - - - - -
    - +
    + + + Serbian + 1 + 45K + + + IE, Slavic + +
    -

    Serbian treebanks

    -
    - - - -
    + +
    ParCoLab is a treebank of Serbian based on literary texts. It was originally developed between 2014 and 2018 (original corpus available here). @@ -4391,11 +3942,10 @@

    Serbian treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4405,35 +3955,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Seri + 1 + - + + + Hokan, Seri + +
    -

    Seri treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4455,11 +3999,10 @@

    Seri treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4469,35 +4012,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Shipibo Konibo + 1 + - + + + Pano-Tacanan + +
    -

    Shipibo Konibo treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4519,11 +4056,10 @@

    Shipibo Konibo treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4533,35 +4069,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Sindhi + 1 + 6K + + + IE, Indic + +
    -

    Sindhi treebanks

    -
    - - - -
    + +
    The Sindhi Universal Dependency Treebank was automatically converted from Sindhi Dependency Treebank (SDTB) which is part of an ongoing effort of creating multi-layered treebanks for Sindhi. @@ -4583,11 +4113,10 @@

    Sindhi treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4597,35 +4126,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Spanish + 1 + <1K + + + IE, Romance + +
    -

    Spanish treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4647,11 +4170,10 @@

    Spanish treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Spanish treebanks. @@ -4663,35 +4185,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Spanish English + 1 + - + + + Code switching + +
    -

    Spanish English treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4713,11 +4229,10 @@

    Spanish English treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4727,35 +4242,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Swahili + 1 + - + + + Niger-Congo, Bantoid + +
    -

    Swahili treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -4777,11 +4286,10 @@

    Swahili treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4791,35 +4299,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Swedish + 1 + - + + + IE, Germanic + +
    -

    Swedish treebanks

    -
    - - - -
    + +
    The Swedish-Eukalyptus treebank has been converted from Eukalyptus, a phrase-structure treebank of contemporary written Swedish. As of now, the conversion has not yet been finished, and no manual corrections have been done. @@ -4841,11 +4343,10 @@

    Swedish treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Swedish treebanks. @@ -4857,35 +4358,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Tagabawa + 1 + <1K + + + Austronesian, Greater Central Philippine + +
    -

    Tagabawa treebanks

    -
    - - - -
    + +
    UD_Tagabawa_GJA is a collection of annotated Bagobo-Tagabawa sentences taken from different sources. It is currently under development. @@ -4907,11 +4402,10 @@

    Tagabawa treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -4921,35 +4415,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Tagalog + 1 + 340K + + + Austronesian, Greater Central Philippine + +
    -

    Tagalog treebanks

    -
    - - - -
    + +
    The Tagalog Universal Dependencies NewsCrawl dataset consists of annotated text extracted from the Leipzig Tagalog Corpus. Data included in the Leipzig Tagalog Corpus were crawled from Tagalog-language online news sites. @@ -4972,11 +4460,10 @@

    Tagalog treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Tagalog treebanks. @@ -4988,35 +4475,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Tetun + 1 + - + + + Austronesian, Malayo-Polynesian + +
    -

    Tetun treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -5038,11 +4519,10 @@

    Tetun treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5052,35 +4532,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Thai + 1 + - + + + Tai-Kadai + +
    -

    Thai treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -5102,11 +4576,10 @@

    Thai treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Thai treebanks. @@ -5118,35 +4591,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Tigrinya + 1 + - + + + Afro-Asiatic, Semitic + +
    -

    Tigrinya treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/contributing/release_checklist.html#the-readme-file) for README guidelines) ... @@ -5168,11 +4635,10 @@

    Tigrinya treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5182,35 +4648,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Tunisian Arabic + 1 + - + + + Afro-Asiatic, Semitic + +
    -

    Tunisian Arabic treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -5232,11 +4692,10 @@

    Tunisian Arabic treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5246,35 +4705,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Turkish + 1 + 142K + + + Turkic, Southwestern + +
    -

    Turkish treebanks

    -
    - - - -
    + +
    The UD_Turkish-ULU Treebank, is an automatic conversion of the ULU Treebank @@ -5296,11 +4749,10 @@

    Turkish treebanks

  • Download
  • -

     

    - - -
    +
    + + See here for comparative statistics of Turkish treebanks. @@ -5312,35 +4764,29 @@

    Language documentation

    See the language documentation page. -
    - - - + + + - - - - -
    - +
    + + + Turkmen + 1 + 47K + + + Turkic, Southwestern + +
    -

    Turkmen treebanks

    -
    - - - -
    + +
    UD_Turkmen-TUD is a silver-standard Universal Dependencies treebank for Turkmen, created by translating Turkish UD treebank data into Turkmen and applying silver annotation through cross-lingual transfer. @@ -5362,11 +4808,10 @@

    Turkmen treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5376,35 +4821,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - - - -
    - +
    + + + Tuwari + 1 + - + + + Sepik + +
    -

    Tuwari treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -5426,11 +4865,10 @@

    Tuwari treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5440,35 +4878,29 @@

    Language documentation

    The language hub documentation has not yet been created. -
    - - - + + + - - +
    + + + Uspanteko + 1 + 13K + + + Mayan + +
    - -
    - - -

    Uspanteko treebanks

    -
    - - - -
    + +
    The MesoTree Uspanteko UD treebank consists of a token-balanced set of trees drawn dictionary example sentences and spoken narratives. @@ -5490,11 +4922,10 @@

    Uspanteko treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5504,35 +4935,29 @@

    Language documentation

    The language hub documentation has not yet been created. - - - - + + + - - +
    + + + Uyghur + 1 + - + + + Turkic, Southeastern + +
    - -
    - - -

    Uyghur treebanks

    -
    - - - -
    + +
    ... 1-2 sentences (see [release checklist](http://universaldependencies.org/release_checklist.html#the-readme-file) for README guidelines) ... @@ -5554,11 +4979,10 @@

    Uyghur treebanks

  • Download
  • -

     

    - - -
    +
    + + @@ -5568,6 +4992,7 @@

    Language documentation

    See the language documentation page. - - + + + diff --git a/_layouts/home.html b/_layouts/home.html index bf92a0efaf..a356ff048a 100644 --- a/_layouts/home.html +++ b/_layouts/home.html @@ -11,7 +11,6 @@ - @@ -41,32 +40,9 @@ {% endif %}
    - {{ content }}
    - - - -{% include accordion.html %}