arXiv:2601.12243cs.CVcs.AI2026-01

用标签引导精简视频摘要,关键帧更准更省。

Less is More: Label-Guided Summarization of Procedural and Instructional Videos

  • 通过标签锚定关键帧,结合视觉采样与大模型验证
  • 仅用不到5%帧数,保留84%语义内容,提升33%性能
  • 适合手术、教程等流程类视频,摘要更连贯可信

视频摘要能将长视频转化为易读、简洁的表示,尤其在手术训练等高风险领域价值突出。现有方法从基础视觉特征发展到使用预训练视觉-语言模型理解视频语义和时间流,实现更上下文感知的摘要。本文提出三阶段框架PRISM(基于整合语义与多模态分析的流程表征),生成语义根基扎实的视频摘要。PRISM融合自适应视觉采样、标签驱动的关键帧锚定,以及利用大语言模型进行上下文验证,确保所选帧反映有意义的流程转换,同时过滤通用或幻觉内容,实现跨领域流程视频的上下文连贯摘要。我们在教学视频和活动数据集上评估该方法,采用教学视频参考摘要。尽管仅采样不足原始帧数的5%,摘要仍保留84%的语义内容,相比基线最高提升33%。该方法在流程类和领域特定视频任务中具有强泛化能力,兼具语义对齐与精确性。

原文摘要 · Abstract (English)

Video summarization helps turn long videos into clear, concise representations that are easier to review, document, and analyze, especially in high-stakes domains like surgical training. Prior work has progressed from using basic visual features like color, motion, and structural changes to using pre-trained vision-language models that can better understand what's happening in the video (semantics) and capture temporal flow, resulting in more context-aware video summarization. We propose a three-stage framework, PRISM: Procedural Representation via Integrated Semantic and Multimodal analysis, that produces semantically grounded video summaries. PRISM combines adaptive visual sampling, label-driven keyframe anchoring, and contextual validation using a large language model (LLM). Our method ensures that selected frames reflect meaningful and procedural transitions while filtering out generic or hallucinated content, resulting in contextually coherent summaries across both domain-specific and instructional videos. We evaluate our method on instructional and activity datasets, using reference summaries for instructional videos. Despite sampling fewer than 5% of the original frames, our summaries retain 84% semantic content while improving over baselines by as much as 33%. Our approach generalizes across procedural and domain-specific video tasks, achieving strong performance with both semantic alignment and precision.

视频摘要流程视频大模型关键帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。