arXiv:2605.01659cs.CVcs.AI2026-05被引 1

用自监督强化学习实现高效视频摘要,无需人工标注。

TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning

论文配图:TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
图 1 · 摘自论文原文
  • 先自监督学特征,再用信息论奖励优化关键帧选择。
  • 在多个数据集上超越现有无监督方法,接近有监督模型表现。
  • 适合需要低成本、可泛化的视频摘要场景。

视频内容在安防、教育和社交媒体等领域的快速增长,使得高效理解变得愈发关键。视频摘要通过生成简洁且语义丰富的表示来应对这一挑战,但现有方法通常依赖昂贵的人工标注,跨领域泛化能力差,且因复杂架构导致计算成本高。此外,无监督和弱监督方法在捕捉长时序依赖和语义结构方面普遍逊于监督方法。本文提出TRIMMER(基于时序相对信息最大化的多目标高效强化学习),一种新型自监督强化学习框架。TRIMMER分两阶段运行:首先通过自监督学习获取鲁棒表征,然后利用信息论奖励函数指导时空决策。不同于依赖相似度的目标,本方法引入基于熵的度量以捕捉更高阶时序动态与语义多样性,并直接在选中帧索引上计算奖励,提升效率。大量实验表明,TRIMMER在标准基准上达到无监督/自监督方法的最先进水平,同时与领先监督方法相当,凸显其在可扩展、可泛化视频摘要中的有效性。

原文摘要 · Abstract (English)

The rapid growth of video content across domains such as surveillance, education, and social media has made efficient content understanding increasingly critical. Video summarization addresses this challenge by generating concise yet semantically meaningful representations, but existing approaches often rely on expensive manual annotations, struggle to generalize across domains, and incur significant computational costs due to complex architectures. Moreover, unsupervised and weakly supervised methods typically underperform compared to supervised counterparts in capturing long-range temporal dependencies and semantic structure. In this work, we propose TRIMMER (Temporal Relative Information Maximization for Multi-objective Efficient Reinforcement), a novel self-supervised reinforcement learning framework for video summarization. TRIMMER operates in two stages: it first learns robust representations via self-supervised learning and then performs spatio-temporal decision making through reinforcement learning guided by information-theoretic reward functions. Unlike prior approaches that rely on similarity-based objectives, our method introduces entropy-based metrics to capture higher-order temporal dynamics and semantic diversity, while computing rewards directly over selected frame indices to improve computational efficiency. Extensive experiments on standard benchmarks demonstrate that TRIMMER achieves state-of-the-art performance among unsupervised and self-supervised methods, while remaining competitive with leading supervised approaches, highlighting its effectiveness for scalable and generalizable video summarization.

视频摘要自监督学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。