arXiv:2508.16313cs.LGcs.AI2025-08EMNLP

让大模型从错误中学习,提升多模态推理能力

Retrieval Enhanced Feedback via In-context Neural Error-book

  • 构建神经错误本,用三类查询结构化反馈
  • 推理速度更快,节省计算资源,效果更好
  • 适合需要高效多模态推理的系统开发者

大型语言模型(LLM)的进展显著提升了推理能力,其中上下文学习(ICL)成为无需重训练即可适应新任务的关键技术。尽管以往研究侧重于利用正确示例,近期工作强调从错误中学习对性能提升的重要性。然而,现有方法缺乏对错误进行系统分析与缓解的框架,尤其在多模态大语言模型(MLLM)中,视觉与文本输入的融合增加了复杂性。为此,我们提出REFINE:基于上下文神经错误本的检索增强反馈框架,这是一个教师-学生架构,能系统化组织错误并提供针对性反馈。REFINE引入三种系统化查询——目标反馈、检查反馈和路径反馈,通过优先关注相关视觉信息、诊断关键失败点、制定纠正措施来增强多模态推理。与依赖冗余检索的先前方法不同,REFINE优化了结构化反馈的检索,提升了推理效率、降低了令牌使用量与可扩展性开销。实验结果表明,该方法实现了显著提速、计算成本降低,并成功实现泛化,凸显其在增强多模态推理方面的潜力。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have significantly improved reasoning capabilities, with in-context learning (ICL) emerging as a key technique for adaptation without retraining. While previous works have focused on leveraging correct examples, recent research highlights the importance of learning from errors to enhance performance. However, existing methods lack a structured framework for analyzing and mitigating errors, particularly in Multimodal Large Language Models (MLLMs), where integrating visual and textual inputs adds complexity. To address this issue, we propose REFINE: Retrieval-Enhanced Feedback via In-context Neural Error-book, a teacher-student framework that systematically structures errors and provides targeted feedback. REFINE introduces three systematic queries to construct structured feedback -- Feed-Target, Feed-Check, and Feed-Path -- to enhance multimodal reasoning by prioritizing relevant visual information, diagnosing critical failure points, and formulating corrective actions. Unlike prior approaches that rely on redundant retrievals, REFINE optimizes structured feedback retrieval, improving inference efficiency, token usage, and scalability. Our results demonstrate substantial speedup, reduced computational costs, and successful generalization, highlighting REFINE's potential for enhancing multimodal reasoning.

多模态推理错误学习高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。