arXiv:2505.06594cs.CLcs.CV2025-05被引 1

零样本生成剧本式视频摘要,兼顾画面与台词信息。

Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation

  • 自建剧本表示,融合关键画面、对白与角色信息。
  • 生成摘要含20%更多相关视觉内容,仅需75%视频输入。
  • 提出新评估指标MFactSum,精准衡量多模态摘要质量。

视觉语言模型在总结复杂多模态内容(如整季电视剧)时,常难以平衡视觉与文本信息。本文提出一种零样本视频到文本摘要方法,通过自建剧本表示,将关键视频片段、对白和角色信息整合为统一文档。该方法无需额外标注,仅依赖音频、视频和字幕即可同时生成剧本并命名角色。此外,我们指出现有摘要评价指标难以有效评估多模态内容,因此提出MFactSum——一种兼顾视觉与文本模态的多模态评估指标。在SummScreen3D数据集上使用MFactSum评估表明,该方法优于Gemini 1.5等先进VLM,生成摘要包含20%更多相关视觉信息,且仅需75%的视频输入量。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) often struggle to balance visual and textual information when summarizing complex multimodal inputs, such as entire TV show episodes. In this paper, we propose a zero-shot video-to-text summarization approach that builds its own screenplay representation of an episode, effectively integrating key video moments, dialogue, and character information into a unified document. Unlike previous approaches, we simultaneously generate screenplays and name the characters in zero-shot, using only the audio, video, and transcripts as input. Additionally, we highlight that existing summarization metrics can fail to assess the multimodal content in summaries. To address this, we introduce MFactSum, a multimodal metric that evaluates summaries with respect to both vision and text modalities. Using MFactSum, we evaluate our screenplay summaries on the SummScreen3D dataset, demonstrating superiority against state-of-the-art VLMs such as Gemini 1.5 by generating summaries containing 20% more relevant visual information while requiring 75% less of the video as input.

多模态摘要视频生成评估指标零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。