arXiv:2412.00948cs.CL2024-12被引 13

构建首个非洲低资源语言科学问答与真实性评测基准

Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

  • 人工翻译现有英文基准,构建六种非洲语言的科学问答与安全评测数据集
  • 所有模型在非洲语言上表现显著低于英语,大模型间差距明显
  • 适合关注多语言AI公平性、非洲语言NLP研究者使用

当前大语言模型在知识密集型任务和事实准确性上的评估多集中于高资源语言,因低资源语言(LRLs)数据稀缺。本文提出Uhura——一个针对六种类型多样非洲语言的新型评测基准,涵盖两个任务:基于人类翻译的多选科学问题集Uhura-ARC-Easy,以及测试模型在健康、法律、金融、政治等领域的可信度的安全评测集Uhura-TruthfulQA。我们指出在低资源语言中构建高技术内容基准的挑战,并提出缓解策略。评估显示,如GPT-4o和o1-preview等专有模型在非洲语言上的表现优于Claude系列,而开源模型如Meta的LLaMA和Google的Gemma表现更差。所有模型在英语中的表现均优于非洲语言。结果表明,大模型在科学问答方面能力有限,且在非洲语言中更易生成虚假信息。研究强调需持续提升多语言模型在低资源语言环境下的能力,以确保其在现实场景中的安全可靠。我们已开源Uhura基准与平台,以推动低资源语言NLP研究发展。

原文摘要 · Abstract (English)

Evaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often focus on high-resource languages primarily because datasets for low-resource languages (LRLs) are scarce. In this paper, we present Uhura -- a new benchmark that focuses on two tasks in six typologically-diverse African languages, created via human translation of existing English benchmarks. The first dataset, Uhura-ARC-Easy, is composed of multiple-choice science questions. The second, Uhura-TruthfulQA, is a safety benchmark testing the truthfulness of models on topics including health, law, finance, and politics. We highlight the challenges creating benchmarks with highly technical content for LRLs and outline mitigation strategies. Our evaluation reveals a significant performance gap between proprietary models such as GPT-4o and o1-preview, and Claude models, and open-source models like Meta's LLaMA and Google's Gemma. Additionally, all models perform better in English than in African languages. These results indicate that LMs struggle with answering scientific questions and are more prone to generating false claims in low-resource African languages. Our findings underscore the necessity for continuous improvement of multilingual LM capabilities in LRL settings to ensure safe and reliable use in real-world contexts. We open-source the Uhura Benchmark and Uhura Platform to foster further research and development in NLP for LRLs.

多语言模型低资源语言评测基准科学问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。