arXiv:2509.24008cs.CVcs.AI2025-09被引 14

让视频模型像人一样边思考边选帧,动态获取关键信息。

FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

  • 通过强化学习实现边推理边主动选帧,自适应获取视觉证据。
  • 在MLVU和VideoMME上超越现有模型,显著提升视频理解准确率。
  • 无需逐帧标注,适合需要灵活感知的复杂视频任务研究者。

当前视频理解模型依赖固定的帧采样策略,无论问题需求如何,都处理预设的视觉输入,导致难以根据具体任务动态获取视觉证据,影响在需要广泛时间覆盖或精细空间细节的任务表现。本文提出FrameMind,一种基于强化学习的端到端框架,支持在推理过程中通过帧交错思维链(FiCOT)动态请求视觉信息。不同于传统方法,FrameMind采用多轮交替推理与主动视觉感知机制,利用工具根据知识缺口提取特定帧或视频片段。为训练高效的动态采样策略,我们提出动态分辨率帧采样(DRFS),在学习中引入多样化的时空权衡;并设计DRFS-GRPO算法,通过结果奖励进行群体相对策略优化,无需帧级标注。在MLVU和VideoMME等挑战性基准上的大量实验表明,该方法显著优于现有模型,推动了灵活高效视频理解的性能边界。

原文摘要 · Abstract (English)

Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each question. This static approach limits their ability to adaptively gather visual evidence, leading to suboptimal performance on tasks that require either broad temporal coverage or fine-grained spatial detail. In this paper, we introduce FrameMind, an end-to-end framework trained with reinforcement learning that enables models to dynamically request visual information during reasoning through Frame-Interleaved Chain-of-Thought (FiCOT). Unlike traditional approaches, FrameMind operates in multiple turns where the model alternates between textual reasoning and active visual perception, using tools to extract targeted frames or video clips based on identified knowledge gaps. To train effective dynamic sampling policies, we propose Dynamic Resolution Frame Sampling (DRFS), which exposes models to diverse temporal-spatial trade-offs during learning, and DRFS-GRPO, a group-relative policy optimization algorithm that learns from outcome-based rewards without requiring frame-level annotations. Extensive experiments on challenging benchmarks like MLVU and VideoMME demonstrate that our method significantly outperforms existing models, advancing the state of the art in flexible and efficient video understanding.

视频理解强化学习动态采样推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。