arXiv:2503.08335cs.CV2025-03ICCV

用提示词工程理解长视频,减少人工标注依赖。

Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos

  • 结合ASR与OCR提取音视频文本,通过提示词分析长视频内容。
  • 在讲座和新闻类长视频上验证,展现模型理解能力边界。
  • 适合研究多模态视频理解与提示工程的学者参考。

多模态视频理解通常依赖于带有手动标注字幕的视频片段数据集。然而,在教育和新闻领域的长视频(持续数分钟至数小时)场景下,因需具备领域专业知识的标注者,标注成本显著增加。因此亟需自动化解决方案。近期大型语言模型(LLMs)的发展为利用自动语音识别(ASR)和光学字符识别(OCR)技术,从音频和特定帧中提取文本内容,实现对完整视频的高效理解提供了可能。本文构建了一个包含长视频讲座与新闻视频的数据集,提出基线方法并揭示其在该数据集上的局限性,主张探索提示工程以更全面地理解长视频多模态数据。

原文摘要 · Abstract (English)

Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging when dealing with long-form videos, lasting from minutes to hours, in educational and news domains due to the need for more annotators with subject expertise. Hence, there arises a need for automated solutions. Recent advancements in Large Language Models (LLMs) promise to capture concise and informative content that allows the comprehension of entire videos by leveraging Automatic Speech Recognition (ASR) and Optical Character Recognition (OCR) technologies. ASR provides textual content from audio, while OCR extracts textual content from specific frames. This paper introduces a dataset comprising long-form lectures and news videos. We present baseline approaches to understand their limitations on this dataset and advocate for exploring prompt engineering techniques to comprehend long-form multimodal video datasets comprehensively.

长视频理解提示工程多模态LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。