arXiv:2509.14704cs.AI2025-09

用日式谜题测试大模型的推理洞察力与自我评估能力,发现其存在显著认知瓶颈。

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

  • 设计日式儿童谜题集,专测模型的顿悟式重构与元认知评估能力。
  • 人类平均准确率52.9%,顶尖模型仅达17.6%,远低于人类表现。
  • 揭示模型普遍存在‘生成不验证’问题,暴露推理可靠性关键缺陷。

基准测试饱和与训练数据污染日益模糊大语言模型(LLMs)在推理上的真实进步是否源于对常见题型的熟悉。我们引入NazoNazo基准,一个源自日本儿童谜题、可更新且可扩展的评测数据集,专注于检验类顿悟的表征重构与元认知评估过程。该任务聚焦于标准基准中难以察觉的失败模式。我们整理了201道谜题,并在120题子集上建立人类参照基准(n = 126;平均准确率52.9%)。该基准完全开放,低成本可刷新,且受污染风险低。我们在严格无检索、零样本条件下评估38个前沿模型(2023–2025)。在人类对比子集上,非推理模型准确率为7.6%,推理导向模型为17.6%,而人类平均为52.9%,模型间表现差异显著。定性分析显示,模型生成正确中间答案但未采纳为最终答案,即‘验证失败’。在有可用思考日志的模型中,验证失败占错误答案的5%至39%。此发现揭示生成与验证间的元认知断层,为诊断推理可靠性提供实用框架,并指向改进方向:更好的校准、结构化验证与停止机制。

原文摘要 · Abstract (English)

Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems. We introduce the NazoNazo Benchmark, a renewable and extensible evaluation dataset derived from Japanese children's riddles that isolates a specific class of reasoning processes: insight-like representational restructuring and metacognitive evaluation. Rather than modeling reasoning in general, these tasks provide a focused test of failure modes that are difficult to detect in standard benchmarks. We curate 201 riddles and establish a human reference on a 120-item subset (n = 126; mean accuracy 52.9%). The benchmark is fully open, low-cost to refresh, and designed for continual evaluation under reduced contamination risk. We evaluate 38 frontier LLMs (2023-2025) under a strict retrieval-free, zero-shot protocol. On the human-comparison subset, non-reasoning models achieve 7.6% accuracy and reasoning-oriented models reach 17.6%, compared with a human mean of 52.9%, although performance varies substantially across models. Beyond accuracy, qualitative analysis of model-generated thought-logs identifies a distinctive failure mode, which we call verification failure: models generate a correct intermediate candidate but fail to endorse it as their final answer. This dissociation between candidate generation and endorsement reveals a metacognitive bottleneck: across the models with usable thought-logs, verification failures account for between 5% and 39% of a model's incorrect answers. By isolating the gap between generation and verification, this work provides a practical framework for diagnosing reasoning reliability and suggests concrete directions for improvement, including better calibration, structured verification, and stopping mechanisms.

推理能力元认知模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。