arXiv:2509.21100cs.CV2025-09NeurIPS被引 61

让AI看视频时像人一样反复聚焦重点,提升理解能力

VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception

  • 通过迭代感知逐步聚焦关键时空区域,模拟人类注意力
  • 在15个以上任务中平均提升超5%,显著优于基线模型
  • 适合需要精细视频理解的场景,如智能客服、自动驾驶

在多模态大语言模型(MLLM)中激发推理能力对实现人类级感知与理解至关重要。现有方法主要依赖语言模型分析已解析的视觉内容,受限于静态感知阶段。本文提出视觉测试时扩展(VTTS),一种通过推理过程中迭代感知来增强MLLM推理的新方法。VTTS模仿人类分层注意力机制,基于更新的文本预测,逐步优化对高置信度时空区域的关注。具体而言,采用迭代感知(ITP)机制,结合时空监督的强化学习以优化推理。为支持该范式,我们还构建了专用于迭代感知的VTTS-80K数据集。该设计使MLLM可通过增加感知计算量来提升性能。大量实验验证了VTTS在多种任务和基准上的有效性与泛化性。新提出的VideoChat-R1.5模型相较强基线Qwen2.5VL-3B和-7B,在超过15个涵盖视频对话、视频推理与时空感知的任务上实现平均超过5%的提升。

原文摘要 · Abstract (English)

Inducing reasoning in multimodal large language models (MLLMs) is critical for achieving human-level perception and understanding. Existing methods mainly leverage LLM reasoning to analyze parsed visuals, often limited by static perception stages. This paper introduces Visual Test-Time Scaling (VTTS), a novel approach to enhance MLLMs' reasoning via iterative perception during inference. VTTS mimics humans' hierarchical attention by progressively refining focus on high-confidence spatio-temporal regions, guided by updated textual predictions. Specifically, VTTS employs an Iterative Perception (ITP) mechanism, incorporating reinforcement learning with spatio-temporal supervision to optimize reasoning. To support this paradigm, we also present VTTS-80K, a dataset tailored for iterative perception. These designs allows a MLLM to enhance its performance by increasing its perceptual compute. Extensive experiments validate VTTS's effectiveness and generalization across diverse tasks and benchmarks. Our newly introduced Videochat-R1.5 model has achieved remarkable improvements, with an average increase of over 5\%, compared to robust baselines such as Qwen2.5VL-3B and -7B, across more than 15 benchmarks that encompass video conversation, video reasoning, and spatio-temporal perception.

多模态推理视频理解迭代感知强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。