用算法题检测大模型跨语言能力差异,发现主流模型存在持续性短板。
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
- 基于可生成的算法任务构建跨语言评测基准,确保公平可比。
- 多模型实验显示,主流模型在不同语言间表现不一,差距显著。
- 适合评估模型跨语言推理能力,尤其关注多语言AI研发者。
我们提出一组合成的算法任务,用于检测大型语言模型在跨语言能力上的差异。该基准具有可比性,因所有语言均需完成相同底层任务;可扩展性,任务复杂度可调,适配不同能力的模型;可量化性,每项任务均有明确正确性标准;透明性,任务由简单模板生成,便于审查翻译错误。由于聚焦算法任务,性能差异是跨语言差距的充分但非必要指标。通过大量实验,我们证实该基准揭示了多个前沿模型中持续存在的跨语言能力差距。
原文摘要 · Abstract (English)
We introduce a set of synthetic algorithmic tasks to detect cross-lingual gaps in the abilities of large language models. Our benchmark is commensurate across languages, since it requires models to perform the same underlying task in different languages; scalable, since each task can be generated at varying levels of complexity allowing it to be adapted to models with different capabilities; quantifiable, since every task admits an objective notion of correctness; and transparent, since tasks are generated from simple templates that can be readily audited for translation errors. Because our benchmark focuses on algorithmic tasks, differential performance is a sufficient -- but not necessary -- indicator of cross-lingual gaps. Nevertheless, we show through extensive experiments that our benchmark exposes persistent cross-lingual gaps in multiple state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。