用国际语言奥赛题测试大模型推理,发现小模型也能超大模型。
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

- 用未见过的语言奥赛题评估模型,先解规则再推理。
- 14B模型表现超越两倍大的模型,解码策略比规模更重要。
- 自动评分与评委打分一致,但压缩得分范围,弱者被高估13分。
大模型的推理研究多集中于有明确规则的领域,如数学和编程;而语言谜题则相反:求解者需先发现规则,再进行推理。本文提出IOL-AI挑战,基于2026年国际语言奥林匹克个人赛的未见题目,采用自动评分与官方评审团人工评分相结合的方式(使用与人类选手相同的评分标准)。共收到来自46支队伍的731份提交,每队限用单张T4显卡、30分钟计算时间。我们还对15个无约束前沿及开源模型进行了基准测试,其中Claude Opus 4.8获得相当于金牌的评委分数,而我们提交的资源受限系统得分仅位于参赛者后5%。模型能力并非由规模决定:140亿参数的模型优于规模翻倍的模型,性能提升主要来自解码和输出处理策略而非模型容量。此外,自动评分与评委打分排序一致,但压缩评分区间,使弱系统平均得分被高估约13分,强系统得分被低估。分析表明,尽管前沿模型可能对部分问题语言有先验知识,但并未显著提升解题能力,说明语言推理仍是衡量通用推理能力的有效基准。
原文摘要 · Abstract (English)
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。