测试大模型灵活推理能力,发现顶级模型准确率仅60%以上
RiddleBench: A New Generative Reasoning Benchmark for LLMs
- 设计1737道英文谜题,评估逻辑与空间综合推理能力
- 顶级模型如Gemini、Claude4准确率均低于64%,存在严重幻觉
- 适合研究模型鲁棒性与人类级推理的开发者使用
大型语言模型在诸多现有推理基准上表现强劲,但这些基准主要评估量化等结构化技能,难以衡量人类智能核心的灵活、多维度推理能力。这类能力需融合逻辑推理、空间感知与约束满足,而当前评估体系未能有效测量。为此,我们提出RiddleBench,一个包含1,737个挑战性谜题的英语基准,旨在探测这些核心推理能力。对前沿模型的评估显示其存在根本性缺陷:即使顶级专有模型(Gemini 2.5 Pro、o3、Claude 4 Sonnet)准确率也仅分别为60.30%、63.37%和63.16%。分析揭示深层问题,包括幻觉级联(接受其他模型的错误推理)、因强自我确认偏差导致的纠错失败,以及在约束顺序调整或引入无关信息时性能显著下降。RiddleBench可作为诊断工具,亦可用于指导更鲁棒、可靠的语言模型开发。
原文摘要 · Abstract (English)
Large Language Models have demonstrated strong performance on many established reasoning benchmarks. However, these benchmarks primarily evaluate structured skills like quantitative problem-solving, leaving a gap in assessing flexible, multifaceted reasoning abilities that are central to human intelligence. These abilities require integrating logical deduction with spatial awareness and constraint satisfaction, which current evaluations do not measure well. To address this, we introduce RiddleBench, a benchmark of 1,737 challenging puzzles in English designed to probe these core reasoning capabilities. Evaluation of state-of-the-art models on RiddleBench shows fundamental weaknesses. Even top proprietary models like Gemini 2.5 Pro, o3, and Claude 4 Sonnet achieve accuracy just above 60% (60.30%, 63.37%, and 63.16%). Analysis further reveals deep failures, including hallucination cascades (accepting flawed reasoning from other models) and poor self-correction due to a strong self-confirmation bias. Their reasoning is also fragile, with performance degrading significantly when constraints are reordered or irrelevant information is introduced. RiddleBench functions as a diagnostic tool for these issues and as a resource for guiding the development of more robust and reliable language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。