跨语言抽象推理测试揭示大模型在不同语言中的表现差异。
Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping Games
- 设计多语言分组游戏评测模型的抽象推理能力
- 英语模态下模型表现显著优于其他语言
- 适合关注大模型语言偏见与公平性的研究者
大型语言模型可能因语言模态差异而表现出推理能力偏差,即在某些语言任务中表现优于其他语言,即使内容相似。现有评估多依赖常识或数学等有明确解法的任务,但日常生活中更关键的是脱离固定模式的抽象推理能力。本文受《纽约时报》谜题游戏Connections: GlobalGroup启发,构建了涵盖英语、西班牙语、中文、印地语和阿拉伯语五种语言的分组游戏基准,每种语言提供原生版本与英文翻译版用于对比。我们引入游戏难度度量标准,实现对难度相近任务的可控比较,更准确评估模型表现。实验发现,英语模态下的模型在该抽象推理任务中普遍表现更好,且开源与闭源模型间存在性能差距。
原文摘要 · Abstract (English)
Large language models (LLMs) can exhibit biases in reasoning capabilities due to linguistic modality, performing better on tasks in one language versus another, even with similar content. Most previous works evaluate this through reasoning tasks where reliance on strategies or knowledge can ensure success, such as in commonsense or math tasks. However, abstract reasoning is vital to reasoning for everyday life, where people apply "out-of-the-box thinking" to identify and use patterns for solutions, without a reliance on formulaic approaches. Comparatively, little work has evaluated linguistic biases in this task type. In this paper, we propose a task inspired by the New York Times Connections: GlobalGroup, that evaluates models in an abstract reasoning task across several languages. We constructed a game benchmark with five linguistic backgrounds -- English, Spanish, Chinese, Hindi, and Arabic -- in both the native language and an English translation for comparison. We also proposed game difficulty measurements to evaluate models on games with similar difficulty, enabling a more controlled comparison, which is particularly important in reasoning evaluations. Through experimentation, we find English modalities largely lead to better performance in this abstract reasoning task, and performance disparities between open- and closed-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。