arXiv:2603.26365cs.CV2026-03被引 1

用强化学习动态压缩视频冗余帧,提升效率与精度。

Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning

  • 基于惊喜增强状态的轻量策略网络,捕捉运动与时间动态。
  • 10%保留率下达99.5%性能,预填充速度提升16倍。
  • 适合长视频理解、资源受限场景下的高效模型部署。

多模态大语言模型在视频理解中表现出色,但因视觉标记冗余导致计算成本高昂且出现“上下文退化”问题。现有压缩方法多依赖启发式或固定变换,常与下游任务目标脱节,适应性差。为此,我们提出SCORE(通过强化学习实现惊喜增强的令牌压缩),一个统一的自适应压缩框架。SCORE引入轻量级策略网络,基于包含帧间残差的惊喜增强状态表示,显式建模时间动态与运动显著性。采用分组强化学习优化策略,结合两阶段课程迁移(从伪静态视频到真实动态视频)以稳定训练。在多个视频理解基准上实验表明,SCORE显著优于现有最优方法。值得注意的是,其在10%标记保留率下保持99.5%原始性能,预填充速度提升16倍,为长视频高效理解提供了可扩展解决方案。

原文摘要 · Abstract (English)

Multimodal Large Language Models have demonstrated remarkable capabilities in video understanding, yet face prohibitive computational costs and performance degradation from ''context rot'' due to massive visual token redundancy. Existing compression strategies typically rely on heuristics or fixed transformations that are often decoupled from the downstream task objectives, limiting their adaptability and effectiveness. To address this, we propose SCORE (Surprise-augmented token COmpression via REinforcement learning), a unified framework that learns an adaptive token compression policy. SCORE introduces a lightweight policy network conditioned on a surprise-augmented state representation that incorporates inter-frame residuals to explicitly capture temporal dynamics and motion saliency. We optimize this policy using a group-wise reinforcement learning scheme with a split-advantage estimator, stabilized by a two-stage curriculum transferring from static pseudo-videos to real dynamic videos. Extensive experiments on diverse video understanding benchmarks demonstrate that SCORE significantly outperforms state-of-the-art baselines. Notably, SCORE achieves a 16x prefill speedup while preserving 99.5% of original performance at a 10% retention ratio, offering a scalable solution for efficient long-form video understanding.

视频理解强化学习令牌压缩高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。