让AI通过自我反思提升视觉理解能力,减少幻觉。
Perception in Reflection
- 采用策略与批判双模型交替迭代,实现视觉感知的逐步优化。
- 显著提升图像理解、描述精度,并降低幻觉生成率。
- 适合需要复杂推理与多步操作的智能体应用。
我们提出一种「感知反射」范式,旨在突破当前大视觉语言模型(LVLMs)在初始感知阶段表现不佳的局限。为此,我们设计了反射感知(RePer)机制,通过策略与批判模型的系统性交替,实现视觉感知的迭代优化。该框架基于反射感知学习(RPL),借助精心构建的视觉反射数据集和反射反似然训练,强化模型内在的反思能力。全面实验表明,RePer在图像理解、描述精度及幻觉减少方面均有可量化提升。值得注意的是,模型注意力模式与人类视觉焦点高度对齐,且RPL实现了细粒度、自由形式的偏好对齐。这些进展确立了「感知反射」作为未来多模态智能体的重要范式,尤其适用于需要复杂推理与多步操作的任务。
原文摘要 · Abstract (English)
We present a perception in reflection paradigm designed to transcend the limitations of current large vision-language models (LVLMs), which are expected yet often fail to achieve perfect perception initially. Specifically, we propose Reflective Perception (RePer), a dual-model reflection mechanism that systematically alternates between policy and critic models, enables iterative refinement of visual perception. This framework is powered by Reflective Perceptual Learning (RPL), which reinforces intrinsic reflective capabilities through a methodically constructed visual reflection dataset and reflective unlikelihood training. Comprehensive experimental evaluation demonstrates RePer's quantifiable improvements in image understanding, captioning precision, and hallucination reduction. Notably, RePer achieves strong alignment between model attention patterns and human visual focus, while RPL optimizes fine-grained and free-form preference alignment. These advancements establish perception in reflection as a robust paradigm for future multimodal agents, particularly in tasks requiring complex reasoning and multi-step manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。