arXiv:2511.13026cs.CV2025-11被引 12

让大模型同时反思视频和文字,提升长视频理解能力。

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

  • 设计多模态反思框架,同步分析视频与文本信息
  • 在4个基准上显著提升长视频理解性能
  • 无需额外训练或外部模型,适合通用长视频任务

仅依赖文本反思的自省机制在多数多模态任务中表现良好,但在长视频理解场景中存在明显局限。根本原因在于:(1) 长视频包含更丰富动态的视觉信息,仅反思文本不足以捕捉关键内容;(2) 纯文本反思缺乏跨模态交互,无法在反思过程中充分整合视觉信息。为此,我们提出REVISOR(REflective VIsual Segment Oriented Reasoning)——一种工具增强的多模态自省推理框架。该框架使多模态大模型能在文本与视觉模态间协同构建内省推理过程,显著提升长视频理解能力。为确保强化学习中模型能准确聚焦于与问题相关的视频片段,我们设计了双归因解耦奖励机制(DADR),并集成至GRPO训练策略中,强制模型推理与所选视频证据之间的因果对齐。REVISOR无需额外监督微调或外部模型,在VideoMME、LongVideoBench、MLVU和LVBench四个基准上均取得优异表现。

原文摘要 · Abstract (English)

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitations. The fundamental reasons for this lie in two points: (1)long-form video understanding involves richer and more dynamic visual input, meaning rethinking only the text information is insufficient and necessitates a further rethinking process specifically targeting visual information; (2) purely text-based reflection mechanisms lack cross-modal interaction capabilities, preventing them from fully integrating visual information during reflection. Motivated by these insights, we propose REVISOR (REflective VIsual Segment Oriented Reasoning), a novel framework for tool-augmented multimodal reflection. REVISOR enables MLLMs to collaboratively construct introspective reflection processes across textual and visual modalities, significantly enhancing their reasoning capability for long-form video understanding. To ensure that REVISOR can learn to accurately review video segments highly relevant to the question during reinforcement learning, we designed the Dual Attribution Decoupled Reward (DADR) mechanism. Integrated into the GRPO training strategy, this mechanism enforces causal alignment between the model's reasoning and the selected video evidence. Notably, the REVISOR framework significantly enhances long-form video understanding capability of MLLMs without requiring supplementary supervised fine-tuning or external models, achieving impressive results on four benchmarks including VideoMME, LongVideoBench, MLVU, and LVBench.

多模态长视频自省推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。