Constructing a Lexical Resource of Russian Derivational Morphology.

International Conference on Language Resources and Evaluation (LREC)(2022)

引用 0|浏览9
暂无评分
摘要
Words of any language are to some extent related thought the ways they are formed. For instance, the verb exempl-ify and the noun example-s are both based on the word example, but the verb is derived from it, while the noun is inflected. In Natural Language Processing of Russian, the inflection is satisfactorily processed; however, there are only a few machine-trackable resources that capture derivations even though Russian has both of these morphological processes very rich. Therefore, we devote this paper to improving one of the methods of constructing such resources and to the application of the method to a Russian lexicon, which results in the creation of the largest lexical resource of Russian derivational relations. The resulting database dubbed DeriNet.RU includes more than 300 thousand lexemes connected with more than 164 thousand binary derivational relations. To create such data, we combined the existing machine-learning methods that we improved to manage this goal. The whole approach is evaluated on our newly created data set of manual, parallel annotation. The resulting DeriNet.RU is freely available under an open license agreement.
更多
查看译文
关键词
Russian, language resource, derivational network, derivational morphology, machine learning
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要