arXiv:2503.23679cs.CV2025-03

通过显式建模场景内容分布,提升零样本视频字幕生成质量

The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning

  • 构建名词短语、场景图、完整句子三类记忆库,分层生成提示
  • 在MSR-VTT等三个数据集上CIDEr指标提升3.4%至16.2%
  • 适合需要高精度、完整描述的视频理解任务

零样本视频字幕生成要求模型在无标注视频-文本对训练的情况下生成高质量字幕。现有方法多利用CLIP提取视觉相关文本提示以引导语言模型生成,但往往只关注场景某一要素,忽略其余视觉信息。为解决此问题并生成更准确、完整的字幕,本文提出一种新的渐进式多粒度文本提示策略。该方法构建三类记忆库:名词短语、名词短语的场景图、完整句子。同时引入类别感知检索机制,显式建模特定主题周围自然语言的分布。大量实验表明,该方法在MSR-VTT、MSVD和VATEX基准上的主要指标CIDEr分别提升5.7%、16.2%和3.4%,优于现有最先进方法。

原文摘要 · Abstract (English)

Zero-shot video captioning requires that a model generate high-quality captions without human-annotated video-text pairs for training. State-of-the-art approaches to the problem leverage CLIP to extract visual-relevant textual prompts to guide language models in generating captions. These methods tend to focus on one key aspect of the scene and build a caption that ignores the rest of the visual input. To address this issue, and generate more accurate and complete captions, we propose a novel progressive multi-granularity textual prompting strategy for zero-shot video captioning. Our approach constructs three distinct memory banks, encompassing noun phrases, scene graphs of noun phrases, and entire sentences. Moreover, we introduce a category-aware retrieval mechanism that models the distribution of natural language surrounding the specific topics in question. Extensive experiments demonstrate the effectiveness of our method with 5.7%, 16.2%, and 3.4% improvements in terms of the main metric CIDEr on MSR-VTT, MSVD, and VATEX benchmarks compared to existing state-of-the-art.

视频生成零样本文本提示场景建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。