arXiv:2506.23714cs.CVcs.CL2025-06中稿 · HHAI WS 2025: Work…被引 2

融合文本、音频和面部线索,自动生成更精准的视频摘要

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization

  • 多模态融合提取语义与情感关键帧
  • 引入强调词提升摘要相关性,ROUGE-1达0.7929
  • 适合教育与社交视频内容分析场景

教育、职场及社交领域视频内容激增,亟需超越传统单模态的摘要技术。本文提出一种行为感知的多模态视频摘要框架,整合文本、音频与视觉线索,生成时间对齐的摘要。通过提取韵律特征、文本线索与视觉指标,识别语义与情感重要时刻。关键贡献在于发现跨模态强调词(bonus words),显著提升摘要的语义相关性与表达清晰度。在基于大模型生成伪真实标签(pGT)的评估中,相比经典提取方法Edmundson,文本评价指标提升明显:ROUGE-1从0.4769升至0.7929,BERTScore从0.9152增至0.9536;视频评价指标方面,F1-Score提升近23%。结果表明,多模态融合能生成更全面且具行为洞察力的视频摘要。

原文摘要 · Abstract (English)

The increasing volume of video content in educational, professional, and social domains necessitates effective summarization techniques that go beyond traditional unimodal approaches. This paper proposes a behaviour-aware multimodal video summarization framework that integrates textual, audio, and visual cues to generate timestamp-aligned summaries. By extracting prosodic features, textual cues and visual indicators, the framework identifies semantically and emotionally important moments. A key contribution is the identification of bonus words, which are terms emphasized across multiple modalities and used to improve the semantic relevance and expressive clarity of the summaries. The approach is evaluated against pseudo-ground truth (pGT) summaries generated using LLM-based extractive method. Experimental results demonstrate significant improvements over traditional extractive method, such as the Edmundson method, in both text and video-based evaluation metrics. Text-based metrics show ROUGE-1 increasing from 0.4769 to 0.7929 and BERTScore from 0.9152 to 0.9536, while in video-based evaluation, our proposed framework improves F1-Score by almost 23%. The findings underscore the potential of multimodal integration in producing comprehensive and behaviourally informed video summaries.

视频摘要多模态行为感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。