arXiv:2412.12173cs.CLcs.AI2024-12被引 2

用迭代推理提升大模型逻辑能力,实测效果显著

A NotSo Simple Way to Beat Simple Bench

  • 设计多步提示+全局一致性检查,动态优化推理过程
  • 在SimpleBench上,平均准确率提升至AVG@5=89.2%,EAG@5达76.5%
  • 适合关注模型推理鲁棒性与评估框架改进的研究者

本文提出一种新型框架,通过迭代推理与反馈驱动方法增强大语言模型(LLM)的推理能力。针对SimpleBench基准中暴露的逻辑连贯性与现实推理缺陷,我们采用多步提示策略结合全局一致性检查,在Claude 3 Opus、Claude 3.5、GPT-4o和o1-preview等主流模型上进行对比分析。结果表明,迭代推理显著提升性能,标准准确率(AVG@5)达到89.2%,新引入的极端平均指标(EAG@5)达76.5%。分析显示,Claude系列在逻辑一致性上表现优异,而GPT-4o具探索性创造力但对模糊提示敏感。通过案例研究揭示空间与时间推理短板,强调结构化推理框架对突破预训练限制的重要性。本研究为集成动态反馈、自适应重启动与多样化评估指标提供基础,推动模型在复杂多领域问题中的推理进步。

原文摘要 · Abstract (English)

This paper presents a novel framework for enhancing reasoning capabilities in large language models (LLMs) by leveraging iterative reasoning and feedback-driven methodologies. Building on the limitations identified in the SimpleBench benchmark, a dataset designed to evaluate logical coherence and real-world reasoning, we propose a multi-step prompting strategy coupled with global consistency checks to improve model accuracy and robustness. Through comparative analysis of state-of-the-art models, including Claude 3 Opus, Claude 3.5, GPT- 4o, and o1-preview, we demonstrate that iterative reasoning significantly enhances model performance, with improvements observed in both standard accuracy metrics (AVG@5) and a newly introduced metric, Extreme Averaging (EAG@5). Our results reveal model-specific strengths: Claude excels in maintaining logical consistency, while GPT-4o exhibits exploratory creativity but struggles with ambiguous prompts. By analyzing case studies and identifying gaps in spatial and temporal reasoning, we highlight areas for further refinement. The findings underscore the potential of structured reasoning frameworks to address inherent model limitations, irrespective of pretraining methodologies. This study lays the groundwork for integrating dynamic feedback mechanisms, adaptive restart strategies, and diverse evaluation metrics to advance LLM reasoning capabilities across complex and multi-domain problem spaces.

大模型推理迭代推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。