arXiv:2410.10303cs.CL2024-10中稿 · NeurIPS被引 4

对比15种语言大模型翻译偏差,发现低资源语言准确率更低

A Comparative Study of Translation Bias and Accuracy in Multilingual Large Language Models for Cross-Language Claim Verification

  • 分预翻译与自翻译两种方式评估跨语言事实核查效果
  • 低资源语言直接推理准确率显著偏低,大模型自翻译提升效果
  • 适合关注多语种事实核查公平性的研究者与应用开发者

数字虚假信息的兴起促使人们关注多语言大模型在事实核查中的应用。本研究系统评估了15种来自五个语系(罗曼语、斯拉夫语、突厥语、印欧语、高加索语)的语言中大模型进行跨语言事实核查时的翻译偏差与准确性。基于XFACT数据集,考察了预翻译与自翻译两种不同翻译方法的影响。以mBERT在英语数据集上的表现作为基准,比较各语言的特定准确率。结果表明,由于训练数据中代表性不足,低资源语言在直接推理中准确率显著降低。此外,更大的模型在自翻译中表现更优,提升了翻译准确性并减少了偏差。研究强调需在低资源语言中实现均衡的多语言训练,以促进可靠事实核查工具的公平获取,并降低不同语言环境下虚假信息传播的风险。

原文摘要 · Abstract (English)

The rise of digital misinformation has heightened interest in using multilingual Large Language Models (LLMs) for fact-checking. This study systematically evaluates translation bias and the effectiveness of LLMs for cross-lingual claim verification across 15 languages from five language families: Romance, Slavic, Turkic, Indo-Aryan, and Kartvelian. Using the XFACT dataset to assess their impact on accuracy and bias, we investigate two distinct translation methods: pre-translation and self-translation. We use mBERT's performance on the English dataset as a baseline to compare language-specific accuracies. Our findings reveal that low-resource languages exhibit significantly lower accuracy in direct inference due to underrepresentation in the training data. Furthermore, larger models demonstrate superior performance in self-translation, improving translation accuracy and reducing bias. These results highlight the need for balanced multilingual training, especially in low-resource languages, to promote equitable access to reliable fact-checking tools and minimize the risk of spreading misinformation in different linguistic contexts.

多语言模型事实核查翻译偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。