arXiv:2509.14752cs.CL2025-09被引 1

KAIO是首个能评估前沿大模型的韩语数学推理基准,推动韩语AI进步。

KAIO: A Collection of More Challenging Korean Questions

  • 构建面向韩语数学推理的长链逻辑题集,避免现有基准过时
  • 当前最强模型GPT-5仅达62.8分,开源模型普遍低于30分,仍有巨大提升空间
  • 私有发布防污染,待顶尖模型超80%准确后再公开,持续迭代

随着中/后训练技术的发展,大语言模型能力快速提升,但传统评测基准(如MMLU、GPQA-D)很快饱和,难以追踪前沿进展。这一问题在韩语领域尤为突出:现有韩语基准数量少、多为翻译或范围狭窄,更新慢,导致过早饱和与数据污染。为此,我们推出KAIO——首个以数学为中心、强调长链推理的韩语基准。相较于已接近饱和的现有韩语评测集,KAIO仍具显著未饱和性:当前最佳模型GPT-5得分为62.8,次优的Gemini-2.5-Pro为52.3,而开源模型如Qwen3-235B和DeepSeek-R1均低于30分,显示充足进步空间,可有效追踪韩语领域前沿模型发展。为防止污染,KAIO将保持私有,通过隔离评估器提供,直至公开模型达到至少80%准确率后才释放数据集,并推出更难版本进行迭代。

原文摘要 · Abstract (English)

With the advancement of mid/post-training techniques, LLMs are pushing their boundaries at an accelerated pace. Legacy benchmarks saturate quickly (e.g., broad suites like MMLU over the years, newer ones like GPQA-D even faster), which makes frontier progress hard to track. The problem is especially acute in Korean: widely used benchmarks are fewer, often translated or narrow in scope, and updated more slowly, so saturation and contamination arrive sooner. Accordingly, at this moment, there is no Korean benchmark capable of evaluating and ranking frontier models. To bridge this gap, we introduce KAIO, a Korean, math-centric benchmark that stresses long-chain reasoning. Unlike recent Korean suites that are at or near saturation, KAIO remains far from saturated: the best-performing model, GPT-5, attains 62.8, followed by Gemini-2.5-Pro (52.3). Open models such as Qwen3-235B and DeepSeek-R1 cluster falls below 30, demonstrating substantial headroom, enabling robust tracking of frontier progress in Korean. To reduce contamination, KAIO will remain private and be served via a held-out evaluator until the best publicly known model reaches at least 80% accuracy, after which we will release the set and iterate to a harder version.

韩语评测数学推理长链思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。