arXiv:2509.00325cs.CLcs.IR2025-09

让大模型自己发现推理漏洞并改进,提升回答质量。

GIER: Gap-Driven Self-Refinement for Large Language Models

  • 用自然语言描述推理缺陷,驱动模型自我反思与修正。
  • 在三个任务中显著提升推理质量与依据准确性,不损失任务正确率。
  • 适合需要高质量推理的场景,如科研、法律分析等。

我们提出GIER(Gap-driven Iterative Enhancement of Responses),一种通过自省与迭代修正来提升大语言模型输出的通用框架,基于概念性质量标准进行优化。不同于依赖示范、例子或思维链模板的提示策略,GIER利用对推理缺口的自然语言描述,引导模型反复批判并改进自身输出,以更好地满足这些标准。在三个推理密集型任务(SciFact、PrivacyQA、e-SNLI)和四种大模型(GPT-4.1、GPT-4o Mini、Gemini 1.5 Pro、Llama 3.3 70B)上,GIER在不降低任务准确率的前提下,显著提升了推理合理性、依据性和推理一致性。分析表明,模型不仅能理解抽象的概念性缺口,还能将其转化为具体的推理改进。

原文摘要 · Abstract (English)

We introduce GIER (Gap-driven Iterative Enhancement of Responses), a general framework for improving large language model (LLM) outputs through self-reflection and revision based on conceptual quality criteria. Unlike prompting strategies that rely on demonstrations, examples, or chain-of-thought templates, GIER utilizes natural language descriptions of reasoning gaps, and prompts a model to iteratively critique and refine its own outputs to better satisfy these criteria. Across three reasoning-intensive tasks (SciFact, PrivacyQA, and e-SNLI) and four LLMs (GPT-4.1, GPT-4o Mini, Gemini 1.5 Pro, and Llama 3.3 70B), GIER improves rationale quality, grounding, and reasoning alignment without degrading task accuracy. Our analysis demonstrates that models can not only interpret abstract conceptual gaps but also translate them into concrete reasoning improvements.

大模型优化自反思推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。