Developing Lrs For Non-Scheduled Indian Languages A Case Of Magahi

LTC(2014)

引用 0|浏览2
暂无评分
摘要
Magahi is an Indo-Aryan Language, spoken mainly in the Eastern parts of India. Despite having a significant number of speakers, there has been virtually no language resource (LR) or language technology (LT) developed for the language, mainly because of its status as a non-scheduled language. The present paper describes an attempt to develop an annotated corpus of Magahi. The data is mainly taken from a couple of blogs in Magahi, some collection of stories in Magahi and the recordings of conversation in Magahi and it is annotated at the POS level using BIS tagset.
更多
查看译文
关键词
LRL,Magahi,Magahi corpus,Magahi annotation,Non-scheduled languages,Annotated corpora,Language resources,ILCIANN
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要