arXiv:2506.20588cs.CV2025-06被引 3

无需标注,用新损失函数实现高效视频摘要,性能媲美监督模型。

TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness

  • 基于马尔可夫过程设计损失函数,无须注意力机制
  • 在SUMME和TVSUM上超越所有无监督方法,接近有监督模型
  • 适合追求高效、泛化强的视频摘要研究者

视频内容日益普及,高效获取关键信息的需求推动了视频摘要与视频亮点提取的研究。然而,现有先进方法多依赖监督标注或计算昂贵的注意力模型,对分布偏移敏感,跨数据集泛化能力差。本文提出一种开创性的自监督视频摘要框架,无需注意力、RNN或Transformer,即可同时捕捉空间与时间依赖关系。该框架融合新型马尔可夫过程驱动的损失函数与两阶段自监督学习范式,兼顾性能与效率。在SUMME和TVSUM数据集上达到当前最优无监督表现,且媲美最佳有监督模型,展示了无需标注的高效架构潜力,为更通用的视频摘要技术铺路,并挑战了对复杂架构的过度依赖。

原文摘要 · Abstract (English)

The increasing ubiquity of video content and the corresponding demand for efficient access to meaningful information have elevated video summarization and video highlights as a vital research area. However, many state-of-the-art methods depend heavily either on supervised annotations or on attention-based models, which are computationally expensive and brittle in the face of distribution shifts that hinder cross-domain applicability across datasets. We introduce a pioneering self-supervised video summarization model that captures both spatial and temporal dependencies without the overhead of attention, RNNs, or transformers. Our framework integrates a novel set of Markov process-driven loss metrics and a two-stage self supervised learning paradigm that ensures both performance and efficiency. Our approach achieves state-of-the-art performance on the SUMME and TVSUM datasets, outperforming all existing unsupervised methods. It also rivals the best supervised models, demonstrating the potential for efficient, annotation-free architectures. This paves the way for more generalizable video summarization techniques and challenges the prevailing reliance on complex architectures.

视频摘要自监督高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。