谷歌浏览器插件
订阅小程序
在清言上使用

Ulysses Tesemõ: a New Large Corpus for Brazilian Legal and Governmental Domain

Felipe A. Siqueira,Douglas Vitório,Ellen Souza, José A. P. Santos,Hidelberg O. Albuquerque, Márcio S. Dias,Nádia F. F. Silva, André C. P. L. F. de Carvalho, Adriano L. I. Oliveira,Carmelo Bastos-Filho

Language Resources and Evaluation(2024)

引用 0|浏览1
暂无评分
摘要
The increasing use of artificial intelligence methods in the legal field has sparked interest in applying Natural Language Processing techniques to handle legal tasks and reduce the workload of these professionals. However, the availability of legal corpora in Portuguese, especially for the Brazilian legal domain, is limited. Existing resources offer some legal data but lack comprehensive coverage. To address this gap, we present Ulysses Tesemõ, a large corpus specifically built for the Brazilian legal domain. The corpus consists of over 3.5 million files, totaling 30.7 GiB of raw text, collected from 159 sources encompassing judicial, legislative, academic, news, and other related data. The data was collected by scraping public information from governmental websites, emphasizing contents generated over the past two decades. We categorized the obtained files into 30 distinct categories, covering various branches of the Brazilian government and different types of texts. The corpus retains the original content with minimal data transformations, addressing the scarcity of Portuguese legal corpora and providing researchers with a valuable resource for advancing in the research area.
更多
查看译文
关键词
Corpus,Legal domain,Governmental domain,Portuguese language,Brazil
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要