arXiv:2505.23268cs.CVcs.AI2025-05被引 2

用视频和字幕联合生成摘要并找亮点,不依赖人工标注。

Unsupervised Transcript-assisted Video Summarization and Highlight Detection

  • 融合视频帧与字幕,通过强化学习生成摘要
  • 相比纯视觉方法,关键片段识别准确率显著提升
  • 无需标注数据,适合大规模视频处理

视频消费是日常生活的重点,但观看完整视频往往费时。为解决此问题,研究者探索了视频摘要与亮点检测技术。尽管已有工作结合视频帧与字幕,或使用强化学习(RL)进行视频摘要与亮点检测,但尚未有研究将双模态信息整合进统一的强化学习框架中。本文提出一种多模态管道,利用视频帧及其对应字幕,通过模态融合机制生成更精简的视频版本,并检测亮点。该管道在强化学习框架中训练,奖励模型生成多样化且具有代表性的摘要,同时确保包含具有意义字幕内容的视频片段。训练过程无监督,可从大规模未标注数据中学习,克服现有标注数据集规模有限的挑战。实验表明,引入字幕信息在视频摘要与亮点检测任务上表现优于仅依赖视觉内容的方法。

原文摘要 · Abstract (English)

Video consumption is a key part of daily life, but watching entire videos can be tedious. To address this, researchers have explored video summarization and highlight detection to identify key video segments. While some works combine video frames and transcripts, and others tackle video summarization and highlight detection using Reinforcement Learning (RL), no existing work, to the best of our knowledge, integrates both modalities within an RL framework. In this paper, we propose a multimodal pipeline that leverages video frames and their corresponding transcripts to generate a more condensed version of the video and detect highlights using a modality fusion mechanism. The pipeline is trained within an RL framework, which rewards the model for generating diverse and representative summaries while ensuring the inclusion of video segments with meaningful transcript content. The unsupervised nature of the training allows for learning from large-scale unannotated datasets, overcoming the challenge posed by the limited size of existing annotated datasets. Our experiments show that using the transcript in video summarization and highlight detection achieves superior results compared to relying solely on the visual content of the video.

视频摘要多模态强化学习无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。