arXiv:2606.27922cs.CVcs.AI2026-06被引 2

通过外部视觉证据实现长视频理解的自纠正,打破模型幻觉循环。

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

论文配图:Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
图 1 · 摘自论文原文
  • 构建三阶段框架:直觉、验证、仲裁,动态检索客观视觉证据
  • 在VideoMME和LongVideoBench上达到顶尖性能,真实纠错率显著提升
  • 提出解耦强化学习算法,解决多阶段推理中的策略耦合问题

当前长视频理解的多模态反思机制主要依赖内部参数的闭环自我反思,缺乏客观外部证据,导致模型常陷入盲目自信并无法修正错误。同时,将强化学习应用于多阶段反思流程会引入严重的策略耦合,且训练数据极度稀缺。为此,本文提出Reflect-R1,首个面向长视频理解的证据驱动自纠正框架。该框架采用三阶段流程:直觉、验证与仲裁。通过动态检索客观视觉证据验证初始直觉,并自主执行多时间搜索以解决冲突,彻底打破幻觉循环。为克服策略耦合,设计了阶段解耦的强化学习算法SD-GRPO,独立计算各推理阶段的优势函数。同时构建包含12万样本的数据集以填补训练数据空白。在VideoMME与LongVideoBench等基准上的大量实验表明,Reflect-R1达到最先进水平,显著提升真实纠错率,实现严格基于客观证据的自纠正。

原文摘要 · Abstract (English)

Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evidence, models are frequently trapped in blind confidence and often fail to correct errors. Furthermore, applying reinforcement learning to multi-stage reflection pipelines introduces severe policy coupling, which is exacerbated by a critical scarcity of dedicated training data. To address these limitations, this work proposes Reflect-R1, the first Evidence-Driven self-correction framework for long video understanding. The framework constructs a three-stage pipeline consisting of intuition, verification, and arbitration. By dynamically retrieving objective visual evidence to verify initial intuitions and autonomously executing multiple temporal searches to resolve conflicts, it completely breaks the hallucination loop. To overcome policy coupling, we design a stage-decoupled reinforcement learning algorithm named SD-GRPO that independently computes advantage functions across different reasoning stages. Concurrently, we construct a dataset of 120K samples to bridge the training data gap. Extensive experiments on benchmarks such as VideoMME and LongVideoBench demonstrate that Reflect-R1 achieves state-of-the-art performance. Our method significantly improves the genuine rectification rate and enables authentic self-correction strictly grounded in objective evidence.

视频理解自纠正证据驱动强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。