Vidi能理解小时级视频,精准定位文本查询的时间段。
Vidi: Large Multimodal Models for Video Understanding and Editing
- 基于多模态输入构建大模型,支持视觉、音频、文本融合理解。
- 可处理长达一小时的视频,对齐文本查询与时间片段,准确率领先现有模型。
- 专为真实编辑场景设计,适合视频剪辑、智能检索等应用开发者。
人类通过连接彼此传递信息,视频已成为互联网上主要的传播与表达媒介。为支持高质量大规模视频内容创作,现代制作流程需同时理解原始素材(如相机拍摄的未剪辑画面)和编辑组件(如视觉特效)。在视频编辑中,模型必须融合多种模态(如视觉、音频、文本),具备丰富背景知识,并处理灵活长度的输入(如时长一小时的原始视频),这对传统模型构成重大挑战。本文介绍Vidi,一个面向广泛视频理解与编辑任务的大型多模态模型家族。首次发布聚焦于时序检索任务,即在输入视频中识别对应给定文本查询的时间区间,该任务在智能编辑中至关重要。模型具备处理长达一小时视频的能力,展现出强大的时间理解能力。为支持真实场景下的全面评估,我们还提出了VUE-TR基准,包含五项关键改进:1)视频时长显著超过现有时序检索数据集;2)支持基于音频的查询;3)查询格式多样,长度不一;4)标注质量高,真实时间范围由人工标注;5)采用优化的交并比(IoU)度量,支持多时间段评估。显著地,Vidi在时序检索任务上超越主流商业模型(如GPT-4o和Gemini),表明其在视频编辑场景中的优势。
原文摘要 · Abstract (English)
Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a modern pipeline requires a comprehensive understanding of both the raw input materials (e.g., the unedited footage captured by cameras) and the editing components (e.g., visual effects). In video editing scenarios, models must process multiple modalities (e.g., vision, audio, text) with strong background knowledge and handle flexible input lengths (e.g., hour-long raw videos), which poses significant challenges for traditional models. In this report, we introduce Vidi, a family of Large Multimodal Models (LMMs) for a wide range of video understand editing scenarios. The first release focuses on temporal retrieval, i.e., identifying the time ranges within the input videos corresponding to a given text query, which plays a critical role in intelligent editing. The model is capable of processing hour-long videos with strong temporal understanding capability, e.g., retrieve time ranges for certain queries. To support a comprehensive evaluation in real-world scenarios, we also present the VUE-TR benchmark, which introduces five key advancements. 1) Video duration: significantly longer than videos of existing temporal retrival datasets, 2) Audio support: includes audio-based queries, 3) Query format: diverse query lengths/formats, 4) Annotation quality: ground-truth time ranges are manually annotated. 5) Evaluation metric: a refined IoU metric to support evaluation over multiple time ranges. Remarkably, Vidi significantly outperforms leading proprietary models, e.g., GPT-4o and Gemini, on the temporal retrieval task, indicating its superiority in video editing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。