用强化学习精炼视频关键帧,提升理解效率与精度
ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
- 通过三阶段机制自动识别语义转折点,动态选择关键帧
- 在VideoMME上达69.1%、LVBench上52.7%,7B模型媲美72B大模型
- 适合需要高效视频理解的场景,如长视频分析与实时应用
当前最先进的视频理解方法面临两大挑战:密集视频内容逐帧处理计算不可行,以及简单均匀采样难以识别语义关键帧。本文提出新型视频理解框架ChronoForge-RL,结合时间顶点蒸馏(TAD)与关键帧感知组相对策略优化(KF-GRPO)。具体而言,设计可微分的关键帧选择机制,通过三阶段流程系统识别语义转折点,在提升计算效率的同时保留时序信息。TAD利用变化评分、转折检测与优先级蒸馏,选取最具信息量的帧;KF-GRPO引入对比学习范式与显著性增强奖励机制,明确激励模型利用帧内容与时序关系。实验显示,ChronoForge-RL在VideoMME上达到69.1%、LVBench上52.7%,显著优于基线方法,且7B参数模型性能接近72B参数模型。
原文摘要 · Abstract (English)
Current state-of-the-art video understanding methods typically struggle with two critical challenges: (1) the computational infeasibility of processing every frame in dense video content and (2) the difficulty in identifying semantically significant frames through naive uniform sampling strategies. In this paper, we propose a novel video understanding framework, called ChronoForge-RL, which combines Temporal Apex Distillation (TAD) and KeyFrame-aware Group Relative Policy Optimization (KF-GRPO) to tackle these issues. Concretely, we introduce a differentiable keyframe selection mechanism that systematically identifies semantic inflection points through a three-stage process to enhance computational efficiency while preserving temporal information. Then, two particular modules are proposed to enable effective temporal reasoning: Firstly, TAD leverages variation scoring, inflection detection, and prioritized distillation to select the most informative frames. Secondly, we introduce KF-GRPO which implements a contrastive learning paradigm with a saliency-enhanced reward mechanism that explicitly incentivizes models to leverage both frame content and temporal relationships. Finally, our proposed ChronoForge-RL achieves 69.1% on VideoMME and 52.7% on LVBench compared to baseline methods, clearly surpassing previous approaches while enabling our 7B parameter model to achieve performance comparable to 72B parameter alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。