arXiv:2512.01077cs.CLcs.HC2025-12中稿 · AACL 2025被引 1

构建10种濒危印地语族语言的食谱数据集,助力文化传承与公平语言技术发展。

ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages

  • 从印度东部偏远社区收集1060个传统食谱,采用低数字素养友好界面。
  • 主流大模型翻译质量差,但提供文化背景和示例后显著提升。
  • 适合关注濒危语言、文化保护与公平AI的研究者使用。

我们提出一个基于文化背景的多模态数据集ELR-1000,包含来自印度东部偏远地区农村社区的1060个传统食谱,涵盖10种濒危语言。这些食谱富含语言与文化细节,通过为低数字素养用户设计的移动端接口采集。该数据集不仅记录烹饪实践,还保留了原住民饮食传统的社会文化语境。我们评估了多个先进大语言模型(LLMs)在将食谱翻译为英文时的表现,发现尽管模型能力强大,但在低资源、文化特异性语言上仍表现不佳。然而,提供针对性上下文——包括语言背景、翻译示例及文化保育指南——显著提升了翻译质量。研究强调需为弱势语言与领域建立基准,以推动公平且具文化敏感性的语言技术发展。作为本工作的一部分,我们向NLP社区开放发布ELR-1000数据集,期望激励面向濒危语言的语言技术开发。

原文摘要 · Abstract (English)

We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered languages. These recipes, rich in linguistic and cultural nuance, were collected using a mobile interface designed for contributors with low digital literacy. Endangered Language Recipes (ELR)-1000 -- captures not only culinary practices but also the socio-cultural context embedded in indigenous food traditions. We evaluate the performance of several state-of-the-art large language models (LLMs) on translating these recipes into English and find the following: despite the models' capabilities, they struggle with low-resource, culturally-specific language. However, we observe that providing targeted context -- including background information about the languages, translation examples, and guidelines for cultural preservation -- leads to significant improvements in translation quality. Our results underscore the need for benchmarks that cater to underrepresented languages and domains to advance equitable and culturally-aware language technologies. As part of this work, we release the ELR-1000 dataset to the NLP community, hoping it motivates the development of language technologies for endangered languages.

濒危语言文化保护多模态数据集公平AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。