arXiv:2602.14488cs.CLcs.AI2026-02

用多模型协作标注构建孟加拉语检索数据集,验证跨语言复用的可靠性风险

BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR

  • 通过多模型一致性校验与人工评估,构建高质量孟加拉语信息检索数据集
  • 发现跨语言迁移中语义保留差异大,部分翻译导致任务有效性下降30%以上
  • 适合低资源语言研究者参考数据构建流程与跨语言复用的风险边界

低资源语言的信息检索(IR)受限于高质量、任务特定标注数据的稀缺。人工标注成本高且难以扩展,而使用大语言模型(LLMs)自动标注又存在标签可靠性、偏见和评估有效性问题。本文提出一种BETA标注框架,采用来自不同模型家族的多个LLM annotator,结合上下文对齐、一致性检查与多数投票机制,最终通过人工评估验证标签质量,构建了首个孟加拉语信息检索数据集。此外,我们进一步检验其他低资源语言数据集经单跳机器翻译后能否有效复用。基于多个语言对的LLM翻译实验表明,语义保留和任务有效性存在显著差异,反映出语言依赖性偏差与不一致的语义传递,直接影响跨语言数据复用的可靠性。该研究揭示了LLM辅助数据构建的潜力与局限,为低资源语言场景下更可靠的基准与评估流程提供了实证依据。

原文摘要 · Abstract (English)

IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensive and difficult to scale, while using large language models (LLMs) as automated annotators introduces concerns about label reliability, bias, and evaluation validity. This work presents a Bangla IR dataset constructed using a BETA-labeling framework involving multiple LLM annotators from diverse model families. The framework incorporates contextual alignment, consistency checks, and majority agreement, followed by human evaluation to verify label quality. Beyond dataset creation, we examine whether IR datasets from other low-resource languages can be effectively reused through one-hop machine translation. Using LLM-based translation across multiple language pairs, we experimented on meaning preservation and task validity between source and translated datasets. Our experiment reveal substantial variation across languages, reflecting language-dependent biases and inconsistent semantic preservation that directly affect the reliability of cross-lingual dataset reuse. Overall, this study highlights both the potential and limitations of LLM-assisted dataset creation for low-resource IR. It provides empirical evidence of the risks associated with cross-lingual dataset reuse and offers practical guidance for constructing more reliable benchmarks and evaluation pipelines in low-resource language settings.

低资源语言数据构建跨语言迁移LLM标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。