让视觉语言模型学会自我纠错,通过自生成数据提升推理能力。
Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks
- 用两轮自纠错生成正负样本,通过直接偏好优化训练模型。
- 在无外部反馈下,模型性能显著提升,避免重复错误。
- 适合研究模型自我改进机制或提升视觉推理质量的学者。
尽管视觉语言模型(VLMs)在视觉与语言推理任务中表现卓越,但其输出常存在缺陷。自我纠错通过指导模型修正自身回答,是解决该问题的可行方案。现有研究多聚焦于大语言模型(LLMs),而对包含视觉与语言信息的VLMs的自我纠错能力仍缺乏探索。本研究考察了VLMs在推理与微调阶段的自我纠错能力。提出自纠错学习(SCL)框架,使VLMs通过自生成的纠错数据,利用直接偏好优化(DPO)实现无需外部反馈的自我提升。具体而言,在推理阶段采用两轮自纠错获取初始与修正回答,基于正确性收集偏好与非偏好样本。实验表明,未经额外微调与外部反馈时,VLMs在迭代推理中难以有效纠错;但通过将自生成的纠错数据分为偏好与非偏好样本进行偏好微调后,模型性能得以提升,并能避免先前错误。研究强调,自我纠错不仅是修正过程,更应通过额外训练增强模型推理能力,使其直接生成高质量回答,无需进一步修正。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) have shown remarkable abilities in visual and language reasoning tasks, they invariably generate flawed responses. Self-correction that instructs models to refine their outputs presents a promising solution to this issue. Previous studies have mainly concentrated on Large Language Models (LLMs), while the self-correction abilities of VLMs, particularly concerning both visual and linguistic information, remain largely unexamined. This study investigates the self-correction capabilities of VLMs during both inference and fine-tuning stages. We introduce a Self-Correction Learning (SCL) approach that enables VLMs to learn from their self-generated self-correction data through Direct Preference Optimization (DPO) without relying on external feedback, facilitating self-improvement. Specifically, we collect preferred and disfavored samples based on the correctness of initial and refined responses, which are obtained by two-turn self-correction with VLMs during the inference stage. Experimental results demonstrate that although VLMs struggle to self-correct effectively during iterative inference without additional fine-tuning and external feedback, they can enhance their performance and avoid previous mistakes through preference fine-tuning when their self-generated self-correction data are categorized into preferred and disfavored samples. This study emphasizes that self-correction is not merely a refinement process; rather, it should enhance the reasoning abilities of models through additional training, enabling them to generate high-quality responses directly without further refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。