arXiv:2511.22173cs.CL2025-11被引 10

测试大模型自我修正能力,发现需外部反馈才能显著提升。

RefineBench: Evaluating Refinement Capability of Language Models via Checklists

  • 用11个领域1000道题构建检查清单评测框架,模拟真实用户反馈场景。
  • 无指导自修正时,顶尖模型准确率仅31%,多数模型迭代后不升反降。
  • 有明确反馈时,大模型5轮内可接近完美,适合评估改进策略的研究者使用。

语言模型能否自我修正回答?这一问题在真实交互中愈发重要,因用户常提出开放式问题并给出不同层次的反馈。现有研究多聚焦可验证任务(如数学竞赛),而忽略复杂反馈场景。为此,我们提出RefineBench,包含11个领域的1000个挑战性问题,并配套检查清单评估框架。评测两种修正模式:(1) 有指导修正(提供自然语言反馈);(2) 自我修正(无外部引导)。结果显示,在自我修正模式下,前沿模型如Gemini 2.5 Pro和GPT-5基准准确率分别为31.3%和29.1%,且多数模型在多轮迭代中未能持续提升(如Gemini-2.5-Pro仅提升+1.8%,DeepSeek-R1下降-0.1%)。而在有指导修正下,专有模型与大开源模型(>70B)均可借助精准反馈在五轮内逼近理想表现。这表明当前大模型仍缺乏有效自修正能力,而RefineBench为追踪进展提供了可靠平台。

原文摘要 · Abstract (English)

Can language models (LMs) self-refine their own responses? This question is increasingly relevant as a wide range of real-world user interactions involve refinement requests. However, prior studies have largely tested LMs' refinement abilities on verifiable tasks such as competition math or symbolic reasoning with simplified scaffolds, whereas users often pose open-ended queries and provide varying degrees of feedback on what they desire. The recent advent of reasoning models that exhibit self-reflection patterns in their chains-of-thought further motivates this question. To analyze this, we introduce RefineBench, a benchmark of 1,000 challenging problems across 11 domains paired with a checklist-based evaluation framework. We evaluate two refinement modes: (1) guided refinement, where an LM is provided natural language feedback, and (2) self-refinement, where LMs attempt to improve without guidance. In the self-refinement setting, even frontier LMs such as Gemini 2.5 Pro and GPT-5 achieve modest baseline scores of 31.3% and 29.1%, respectively, and most models fail to consistently improve across iterations (e.g., Gemini-2.5-Pro gains only +1.8%, while DeepSeek-R1 declines by -0.1%). By contrast, in guided refinement, both proprietary LMs and large open-weight LMs (>70B) can leverage targeted feedback to refine responses to near-perfect levels within five turns. These findings suggest that frontier LMs require breakthroughs to self-refine their incorrect responses, and that RefineBench provides a valuable testbed for tracking progress.

语言模型自修正评测基准AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。