arXiv:2607.23058cs.CLcs.AI2026-07

构建无需翻译的多语言类比推理评估框架,发现模型在非英语语境下表现显著下降。

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

  • 用母语者校验+大模型生成,构建无翻译偏差的多语言类比题
  • 三语测试中模型准确率相比英语下降12至52个百分点
  • 适合关注跨文化推理能力与模型公平性的研究者

多语言推理评估长期依赖将英文基准翻译到其他语言,这引入了语言层面的干扰,无法真正检验基于文化的推理能力。我们提出ADAGE(面向具身评估的类比难度设计基准),一种无需翻译的多语言评估流水线,结合母语者校验与大模型辅助生成,构建具有挑战性的抽象类比推理基准。我们在阿拉伯语、阿姆哈拉语和日语上验证了该方法,评估了14个开源模型。结果发现:在英语谚语推理中表现良好的模型,在三个母语基准上的表现均大幅下降,准确率降低12至52个百分点。我们公开了完整流水线、三个基准数据集及全部评估结果。

原文摘要 · Abstract (English)

Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation), a language-agnostic pipeline that combines native-speaker curation with LLM-assisted generation to construct challenging, translation-free benchmarks for abstract analogical reasoning. We validate ADAGE by constructing benchmarks for Arabic, Amharic, and Japanese. Evaluating 14 open-weight models, we find a consistent cultural reasoning gap: models that perform well on English proverb reasoning struggle substantially on all three native benchmarks, with accuracy dropping by 12--52 percentage points relative to English. We release the pipeline, all three benchmarks, and the full evaluation suite.

类比推理多语言评估文化偏见LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。