评估多语言大模型真实性,发现跨语言差异小于预期。
Truth Knows No Language: Evaluating Truthfulness Beyond English
- 用专业翻译扩展TruthfulQA至4种语言,评估12个开源模型。
- 英语表现最好,巴斯克语最差,但整体差异不大。
- LLM做裁判比选择题更接近人评,通用知识更易跨语言保持。
我们引入了经过专业翻译的TruthfulQA基准扩展版,用于评估巴斯克语、加泰罗尼亚语、加利西亚语和西班牙语中的真实性。目前大语言模型(LLMs)的真实性评估主要集中在英语,而其在多语言环境下的表现仍待深入研究。本研究评估了12个最先进的开源大模型,对比基础模型与指令微调模型,采用人工评估、多项选择指标及LLM作为裁判评分。结果表明,尽管模型在英语中表现最佳,在巴斯克语中最差(资源最少),但跨语言真实性的差异远小于预期。此外,我们发现LLM作为裁判比多项选择指标更贴近人工判断,且信息量在真实性评估中起关键作用。研究还表明,机器翻译可有效扩展真实性基准至其他语言,是专业翻译的可行替代方案。最后,通用知识问题在各语言间处理得更好,而依赖上下文和时间的问题则表现较差,凸显了评估需考虑文化与时间因素。数据集和代码已公开,采用开放许可。
原文摘要 · Abstract (English)
We introduce a professionally translated extension of the TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. Truthfulness evaluations of large language models (LLMs) have primarily been conducted in English. However, the ability of LLMs to maintain truthfulness across languages remains under-explored. Our study evaluates 12 state-of-the-art open LLMs, comparing base and instruction-tuned models using human evaluation, multiple-choice metrics, and LLM-as-a-Judge scoring. Our findings reveal that, while LLMs perform best in English and worst in Basque (the lowest-resourced language), overall truthfulness discrepancies across languages are smaller than anticipated. Furthermore, we show that LLM-as-a-Judge correlates more closely with human judgments than multiple-choice metrics, and that informativeness plays a critical role in truthfulness assessment. Our results also indicate that machine translation provides a viable approach for extending truthfulness benchmarks to additional languages, offering a scalable alternative to professional translation. Finally, we observe that universal knowledge questions are better handled across languages than context- and time-dependent ones, highlighting the need for truthfulness evaluations that account for cultural and temporal variability. Dataset and code are publicly available under open licenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。