arXiv:2511.21086cs.CL2025-11中稿 · LREC 2026

测试大模型在拼写约束下的表现,发现性能差距远超参数规模影响。

Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models

  • 跨模型家族评估39种配置,用58个字谜检验字符级约束满足能力
  • 顶级模型F1达0.761,最低仅0.343,差距达2.2倍,远超参数量提升效果
  • 模型常误判人类高正确率的非常规拼写词,暴露对语义合理性依赖

大语言模型在可控文本生成中需满足严格的拼写约束,但跨模型家族的系统性评估仍有限。本文评估了涵盖三个模型家族(Qwen3、Claude Haiku 4.5、GPT-5-mini)共39种配置,在58个要求字符级约束满足的单词谜题上进行测试。跨家族性能差异显著,最大差距达2.0–2.2倍(F1值从0.761降至0.343),远超同家族内参数量扩展带来的影响(4B到32B仅提升83%)。偏相关分析排除了分词器设计对同家族缩放效应的干扰。思考预算敏感性呈现异质性:高容量模型收益明显(F1提升0.102–0.136),而中等规模模型则趋于饱和或退化,计算资源投入回报不一致。基于每道谜题10,000名人类解题者的难度评分,发现所有家族均存在适度但稳定的校准性(ρ = 0.28–0.42),但在常见但拼写异常的词汇(如“data”、“loll”、“acai”)上系统性失败——人类成功率83–91%,模型误判率达94–98%。这表明模型过度依赖分布合理性,忽视符合约束但拼写非常规的有效模式。

原文摘要 · Abstract (English)

Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains limited. We evaluate 39 configurations spanning three model families (Qwen3, Claude Haiku 4.5, GPT-5-mini) on 58 word puzzles requiring character-level constraint satisfaction. Cross-family differences produce substantially larger performance gaps (2.0-2.2x, F1 = 0.761 vs. 0.343) than parameter scaling within families (83% gain from 4B to 32B scaling), and a partial-correlation analysis rules out tokenizer design as a confound for within-family scaling. Thinking budget sensitivity proves heterogeneous: high-capacity models show strong returns (+0.102 to +0.136 F1), while mid-sized variants saturate or degrade, showing inconsistent compute benefits. Using difficulty ratings from 10,000 human solvers per puzzle, we establish modest but consistent calibration (\r{ho} = 0.28-0.42) across all families, yet identify systematic failures on common words with unusual orthography ("data", "loll", "acai": 83-91% human success, 94-98% model miss rate). These failures point to over-reliance on distributional plausibility that penalizes orthographically atypical but constraint-valid patterns.

语言模型拼写约束人类对齐认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。