评测多语言大模型在中学生问答中的事实准确性,发现其对小语种表现更差。
Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs
- 用中高考水平问题测试Llama3.1多语言回答事实正确性
- 模型在非英语语种中生成更多错误或冗余信息
- 严重加剧对稀有语言的偏见,适合教育科技开发者参考
事实性是教育类工具有效的前提。随着大语言模型在教育领域的应用持续增长,确保各类场景下的正确性至关重要。尽管模型在英语上表现良好,但其在其他语言中的表现仍缺乏充分验证。本文评估了Llama3.1系列模型在回答适合初中和高中学生的事实性问题时的正确性。结果表明,这些模型不仅提供额外且不真实的信息,还加剧了对稀有语言的现有偏见。
原文摘要 · Abstract (English)
Factuality is a necessary precursor to useful educational tools. As adoption of Large Language Models (LLMs) in education continues of grow, ensuring correctness in all settings is paramount. Despite their strong English capabilities, LLM performance in other languages is largely untested. In this work, we evaluate the correctness of the Llama3.1 family of models in answering factual questions appropriate for middle and high school students. We demonstrate that LLMs not only provide extraneous and less truthful information, but also exacerbate existing biases against rare languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。