arXiv:2507.17050cs.CV2025-07ICCV被引 2

无需训练,用大模型自动生成带时间戳的视频描述。

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

  • 用现成多模态大模型做生成、上下文和验证,无需训练。
  • 生成内容更准,时间对齐更好,幻觉减少。
  • 适合需要快速生成视频摘要或问答的场景。

本文提出VideoNarrator,一种无需训练的视频密集描述生成框架,可生成带精确时间戳的结构化视频叙述。尽管多模态大语言模型(MLLMs)在视频理解方面取得进展,但其在时间对齐叙述和陌生场景中仍易产生幻觉。VideoNarrator通过灵活管道,利用现成的MLLMs和视觉-语言模型(VLMs)作为描述生成器、上下文提供者或描述验证者,实现协同增强。实验表明,该机制显著提升叙述质量与准确性,有效减少幻觉并改善时间对齐。此结构化方法不仅增强视频理解,还可用于视频摘要、问答等下游任务,未来或可拓展至广告与营销应用。

原文摘要 · Abstract (English)

In this paper, we introduce VideoNarrator, a novel training-free pipeline designed to generate dense video captions that offer a structured snapshot of video content. These captions offer detailed narrations with precise timestamps, capturing the nuances present in each segment of the video. Despite advancements in multimodal large language models (MLLMs) for video comprehension, these models often struggle with temporally aligned narrations and tend to hallucinate, particularly in unfamiliar scenarios. VideoNarrator addresses these challenges by leveraging a flexible pipeline where off-the-shelf MLLMs and visual-language models (VLMs) can function as caption generators, context providers, or caption verifiers. Our experimental results demonstrate that the synergistic interaction of these components significantly enhances the quality and accuracy of video narrations, effectively reducing hallucinations and improving temporal alignment. This structured approach not only enhances video understanding but also facilitates downstream tasks such as video summarization and video question answering, and can be potentially extended for advertising and marketing applications.

视频生成大模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。