构建可复现的中英韩谜题基准,诊断韩语模型性能差距根源
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

- 通过行为校准难度,生成15类共7500个可复现谜题
- 韩语模型在多数任务上与英语表现相当,但书写系统密集任务差距显著
- 揭示多步音节执行困难是核心瓶颈,适合语言模型评估研究者
我们提出NOLLI,一个程序化生成的英韩谜题基准,用于诊断韩语性能差距的成因。该基准包含15种谜题类型(25个任务;7500个实例),每个实例可种子再生、有唯一解且评分确定。不同于以规模衡量难度,我们通过行为校准:调整生成器使参考模型准确率落在目标区间。其三层设计涵盖匹配直译、基于韩文字母(jamo)的脚本改编,以及根植于韩国文化或正字法的纯韩语任务。评估了15个前沿、开源及韩方开发模型;其中12个整体准确率超过3%,其英韩表现统计等价(±10个百分点内,TOST检验)。书写系统密集任务差距明显:韩语密码比英语低最多68.7个百分点,而相同jamo的数字符号谜题无系统性惩罚;且韩文字母组合准确率能预测韩语密码表现。这些对比具诊断性而非因果性,支持多步子音节执行困难。纯韩语任务分离出规则应用缺陷(方向不一)与亲缘关系缺陷(12个模型均呈正向)。最后,在15类中的7类,从易到难结构大小未增长,表明结构大小无法可靠代表实际难度。
原文摘要 · Abstract (English)
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。