用约束满足问题评估大模型的长链反思推理能力
LR^2Bench: Evaluating Long-chain Reflective Reasoning Capabilities of Large Language Models via Constraint Satisfaction Problems
- 设计六类约束满足任务,测试模型反思与修正能力
- 顶尖模型平均准确率仅20%-23.6%,反映推理短板
- 适合研究大模型逻辑推理与自我修正能力的学者
近期大型推理模型(LRMs)显著提升了大语言模型(LLMs)的推理能力,使其能通过反思机制(如假设生成、回溯、自优化)处理更复杂任务。然而,有效评估此类反思能力仍面临挑战,因缺乏合适基准。为此,我们提出LR$^2$Bench,一个新型基准,用于评估LLMs的长链反思推理能力。该基准包含850个样本,覆盖六类约束满足问题(CSPs),其中反思推理对满足所有约束至关重要。每类任务聚焦不同约束模式,如知识型、逻辑型和空间型约束,全面覆盖多样化求解场景。我们在常规LLMs和LRMs上进行广泛评估,发现即使最先进的模型如DeepSeek-R1和OpenAI o1-preview在该基准上表现不佳,平均精确匹配(Exact Match)得分分别为20.0%和23.6%。结果表明当前LLMs的反思推理能力仍有巨大提升空间。
原文摘要 · Abstract (English)
Recent progress in Large Reasoning Models (LRMs) has significantly enhanced the reasoning abilities of Large Language Models (LLMs), empowering them to tackle increasingly complex tasks through reflection capabilities, such as making assumptions, backtracking, and self-refinement. However, effectively evaluating such reflection capabilities remains challenging due to the lack of appropriate benchmarks. To bridge this gap, we introduce LR$^2$Bench, a novel benchmark designed to evaluate the Long-chain Reflective Reasoning capabilities of LLMs. LR$^2$Bench comprises 850 samples across six Constraint Satisfaction Problems (CSPs) where reflective reasoning is crucial for deriving solutions that meet all given constraints. Each type of task focuses on distinct constraint patterns, such as knowledge-based, logical, and spatial constraints, providing a comprehensive evaluation of diverse problem-solving scenarios. Our extensive evaluation on both conventional LLMs and LRMs reveals that even the most advanced LRMs, such as DeepSeek-R1 and OpenAI o1-preview, struggle with tasks in LR$^2$Bench, achieving an average Exact Match score of only 20.0% and 23.6%, respectively. These findings underscore the significant room for improvement in the reflective reasoning capabilities of current LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。