arXiv:2505.07172cs.CV2025-05

通过添加思维链增强指令微调,有效减少视觉幻觉。

Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning

  • 在指令中加入视觉推理理由,引导模型更准确理解图像
  • 在偏好微调中引入自批判机制,提升响应一致性
  • 适合需要高可靠性多模态推理的应用场景

尽管大型视觉语言模型(LVLMs)在多模态推理任务上取得显著进展,但在解读关联图像时仍易生成与视觉内容不符的回应。人类学习新知识前常通过回顾提纲、总结要点来建立认知框架,而当前指令微调过程缺乏此类准备步骤。本文提出Re-Critic,一种可扩展的推理理由增强框架,将基本规则与链式思维(CoT)作为桥梁,提升模型推理能力。具体而言,Re-Critic构建了视觉推理合成器,自动为原始指令添加推理说明;并通过上下文自批判机制,筛选响应对用于偏好微调,以获取更具上下文一致性的输出。实验表明,使用该增强数据集微调的模型,在减少幻觉任务之外,也提升了更广泛的多模态推理性能。

原文摘要 · Abstract (English)

Despite significant advancements in multimodal reasoning tasks, existing Large Vision-Language Models (LVLMs) are prone to producing visually ungrounded responses when interpreting associated images. In contrast, when humans embark on learning new knowledge, they often rely on a set of fundamental pre-study principles: reviewing outlines to grasp core concepts, summarizing key points to guide their focus and enhance understanding. However, such preparatory actions are notably absent in the current instruction tuning processes. This paper presents Re-Critic, an easily scalable rationale-augmented framework designed to incorporate fundamental rules and chain-of-thought (CoT) as a bridge to enhance reasoning abilities. Specifically, Re-Critic develops a visual rationale synthesizer that scalably augments raw instructions with rationale explanation. To probe more contextually grounded responses, Re-Critic employs an in-context self-critic mechanism to select response pairs for preference tuning. Experiments demonstrate that models fine-tuned with our rationale-augmented dataset yield gains that extend beyond hallucination-specific tasks to broader multimodal reasoning tasks.

视觉推理幻觉抑制指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。