构建首个印度部落语言机器翻译基准,推动数字公平
AdiBhashaa: A Community-Curated Benchmark for Machine Translation into Indian Tribal Languages
- 联合本土使用者共建平行语料库,全程人工校验
- 首次实现比尔语、蒙达里语等四种部落语言的翻译系统
- 强调社区参与与本地知识,适合关注AI公平的研究者
大型语言模型和多语言机器翻译系统正日益影响信息获取,但许多印度部落语言仍被排除在这些技术之外,加剧了教育、治理与数字参与中的结构性不平等。我们提出AdiBhashaa,一个由社区驱动的项目,构建了四种主要印度部落语言——比尔语(Bhili)、蒙达里语(Mundari)、贡迪语(Gondi)和桑塔利语(Santali)——的首个开放平行语料库及基线机器翻译系统。该工作结合参与式数据构建、母语者协作、人机协同验证,并系统评估编码器-解码器模型与大语言模型在这些语言上的表现。除技术成果外,我们还阐明AdiBhashaa如何为更公平的AI研究提供范例:以本地知识为中心,培养边缘化群体早期研究者能力,强调人类验证在语言技术开发中的核心作用。
原文摘要 · Abstract (English)
Large language models and multilingual machine translation (MT) systems increasingly drive access to information, yet many languages of the tribal communities remain effectively invisible in these technologies. This invisibility exacerbates existing structural inequities in education, governance, and digital participation. We present AdiBhashaa, a community-driven initiative that constructs the first open parallel corpora and baseline MT systems for four major Indian tribal languages-Bhili, Mundari, Gondi, and Santali. This work combines participatory data creation with native speakers, human-in-the-loop validation, and systematic evaluation of both encoder-decoder MT models and large language models. In addition to reporting technical findings, we articulate how AdiBhashaa illustrates a possible model for more equitable AI research: it centers local expertise, builds capacity among early-career researchers from marginalized communities, and foregrounds human validation in the development of language technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。