Attribution

LexicRo's /analyze endpoint is built on openly licensed language resources. This file records what they are, how they are used, and what their licences require.


MULTEXT-East Romanian word-form lexicon

Used for: lemma lookup at request time, and to derive the suffix-rule lemmatiser for out-of-vocabulary words. Around 73% of tokens in a typical request are resolved directly from this lexicon.

Erjavec, Tomaž et al. MULTEXT-East free lexicons 4.0, Slovenian language resource repository CLARIN.SI, 2010. http://hdl.handle.net/11356/1041

The MULTEXT-East morpho-lexical resources are described in:

Dan Tufiş, Ide N., Erjavec T. Standardised Specifications, Development and Assessment of Large Morpho-Lexical Resources for Six Central and Eastern European Languages. First International Conference on Language Resources and Evaluation, Granada, 28–30 May 1998, pp. 233–240.

Dimitrova L., Erjavec T., Ide N., Kaalep H.J., Petkevič V., Tufiş D. MULTEXT-East: Parallel and Comparable Corpora and Lexicons for Six Central and Eastern European Languages. COLING-ACL, Montréal, 1998.

The Romanian portion derives from work by Dan Tufiș and colleagues at RACAI (Research Institute for Artificial Intelligence "Mihai Drăgănescu", Romanian Academy). The MULTEXT-East morphosyntactic specification defines the MSD tagset that LexicRo's MSD→UD bridge converts.


UD Romanian RRT (RoRefTrees)

Used for: training and evaluating the morphological tagger and the lemma head. Every accuracy figure LexicRo publishes is measured on this treebank's test split.

Barbu Mititelu, Verginica, Radu Ion, Radu Simionescu, Elena Irimia and Cenel-Augusto Perez. The Romanian Treebank Annotated According to Universal Dependencies. Proceedings of the Tenth International Conference on Natural Language Processing (HrTAL 2016).

Built on RACAI-RoTb (Irimia and Barbu Mititelu, 2015) and UAIC-RoTb (Perez, 2014). Development supported by CNCS-UEFISCDI project PN-II-RU-TE-2014-4-1362 and COST action CA21167 UniDive.

Annotation provenance, per the treebank's own metadata: lemmas and XPOS are automatic; UPOS and features are converted with corrections; dependency relations are manual native. LexicRo's reported lemma accuracy is therefore agreement with an automatic annotation, and the MSD→UD bridge validation is agreement between two converters. This is the standard benchmark for Romanian and the figures are comparable to published work, but the distinction is worth stating rather than glossing.


Romanian BERT

Used for: the contextual encoder that disambiguates readings a lexicon cannot resolve alone.

Dumitrescu, Stefan Daniel, Andrei-Marius Avram and Sampo Pyysalo. The Birth of Romanian BERT. Findings of EMNLP 2020. https://arxiv.org/abs/2009.08712


Universal Dependencies

The UPOS tags and morphological features returned by /analyze follow the Universal Dependencies v2 guidelines.


Scope of use

For clarity about how each CC BY-SA 4.0 resource enters the service:

A contributor to the UD Romanian RRT treebank has confirmed CC BY-SA 4.0 as the applicable licence for that resource. Attribution as required by both licences is given above, and is served publicly at /attribution.


Acknowledgements

Alex Popescu — voroave.ro — for pointing out DEXonline's official dataset dumps at dexonline.ro/tools, along with dexonline.ro/surse and clre.solirom.ro. voroave.ro is his own project: an effort to surface Romanian words that are dusty but not yet archaic.


verbecc

Used for: the /conjugate endpoint's verb conjugation tables. LexicRo reshapes verbecc's output into its own response schema and synthesises the condițional mood that verbecc declares but does not populate. LexicRo also applies three targeted transformations to individual forms: it normalises legacy cedilla characters to Romanian's comma-below diacritics; it serves the infinitiv mood from verbecc's own verb.infinitive field — the looked-up lemma — rather than from the mood's generated output; and it composes the negative imperative's second-person-singular form as nu + infinitive for every verb that has such a form, applying the invariant paradigm, skipping only those verbs where verbecc records the form non-existent. Every other form is verbecc's own, unmodified.

The latter two transformations were introduced to route around corrupted templates in verbecc's Romanian data — the face family's infinitiv and negative imperative rendered as fudrir;odrir, and similar damage to a avea and a vrea. Those templates were corrected upstream in 2.0.3, so both transformations now agree with verbecc's own values rather than replacing them. Both are retained, because each rests on a claim that does not depend on the defect: verb.infinitive is the datum verbecc looked up rather than one it generated, and the negative imperative 2sg is invariantly nu + infinitive in Romanian. Neither substitutes a linguistic judgement for verbecc's.

verbecc is used unmodified, as an installed library dependency. LexicRo does not distribute it, link it statically, or ship a modified version of it.

Every form served by /conjugate carries a source field recording whether it came from verbecc or was derived by LexicRo, and every response carries a notes array recording the source's known limitations. See the /conjugate guide.