多语言大模型在印度语系中更易幻觉,英文提问更准。
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
- 用同一问题对比英、印地语回答,检验模型可靠性。
- 英文提问时准确率更高,印地语中幻觉率显著上升。
- 适合关注低资源语言模型可信度的研究者参考。
多语言大语言模型(LLMs)在英语等高资源语言中表现优异,但在低资源语言(尤其是印地语系语言)中的事实准确性仍待考察。本研究通过比较 GPT-4o、Gemma-2-9B、Gemma-2-2B 与 Llama-3.1-8B 在英语和19种印地语系语言上的表现,使用包含英、印地语问答对的 IndicQuest 数据集,对相同问题分别以英语及对应印地语翻译进行提问,评估模型在不同语言下的事实准确性。结果表明,即使问题源自印地文化背景,模型在英文提问下表现更优;尤其在低资源印地语中,模型更易产生幻觉,暴露出当前多语言模型在跨语言理解上的局限性。
原文摘要 · Abstract (English)
Multilingual Large Language Models (LLMs) have demonstrated significant effectiveness across various languages, particularly in high-resource languages such as English. However, their performance in terms of factual accuracy across other low-resource languages, especially Indic languages, remains an area of investigation. In this study, we assess the factual accuracy of LLMs - GPT-4o, Gemma-2-9B, Gemma-2-2B, and Llama-3.1-8B - by comparing their performance in English and Indic languages using the IndicQuest dataset, which contains question-answer pairs in English and 19 Indic languages. By asking the same questions in English and their respective Indic translations, we analyze whether the models are more reliable for regional context questions in Indic languages or when operating in English. Our findings reveal that LLMs often perform better in English, even for questions rooted in Indic contexts. Notably, we observe a higher tendency for hallucination in responses generated in low-resource Indic languages, highlighting challenges in the multilingual understanding capabilities of current LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。