Vidi2.5提升视频理解与创作能力,支持精细时空定位和剧情推理。
Vidi2.5: Large Multimodal Models for Video Understanding and Creation
- 引入细粒度时空定位与剧情理解新模型,支持文本查询精准定位目标对象
- 在VUE-STG与VUE-PLOT基准上超越Gemini 3 Pro Preview等主流系统
- 适用于视频编辑规划等复杂真实场景,适合多模态推理研究者使用
视频已成为互联网上主要的传播与创作媒介,推动对可扩展、高质量视频生产的需求。Vidi系列模型持续演进,第二版Vidi2在多模态时序检索(TR)任务中达到领先水平,并增强细粒度时空定位(STG)能力,扩展至视频问答(Video QA),实现全面的多模态推理。给定文本查询,Vidi2可识别对应时间戳及目标物体的边界框。为此,我们提出新基准VUE-STG,显著优于现有STG数据集;同时升级原VUE-TR为VUE-TR-V2,实现更均衡的时长与查询分布。令人瞩目的是,Vidi2在两个基准上均显著超越Gemini 3 Pro Preview和GPT-5等领先闭源系统,且在视频QA上表现媲美同规模开源模型。最新发布的Vidi2.5进一步强化了STG能力,微调提升了TR与Video QA性能,并推出Vidi2.5-Think模型用于复杂剧情理解。为全面评估剧情理解,我们提出包含角色与推理双赛道的VUE-PLOT基准。值得注意的是,Vidi2.5-Think在角色理解上优于Gemini 3 Pro Preview,复杂推理性能相当。此外,我们在真实世界视频编辑规划任务中验证了Vidi2.5的有效性。
原文摘要 · Abstract (English)
Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to evolve toward next-generation video creation and have achieved state-of-the-art performance in multimodal temporal retrieval (TR). In its second release, Vidi2 advances video understanding with fine-grained spatio-temporal grounding (STG) and extends its capability to video question answering (Video QA), enabling comprehensive multimodal reasoning. Given a text query, Vidi2 can identify not only the corresponding timestamps but also the bounding boxes of target objects within the output time ranges. To enable comprehensive evaluation of STG, we introduce a new benchmark, VUE-STG, which offers critical improvements over existing STG datasets. In addition, we upgrade the previous VUE-TR benchmark to VUE-TR-V2, achieving a more balanced duration and query distribution. Remarkably, the Vidi2 model substantially outperforms leading proprietary systems, such as Gemini 3 Pro Preview and GPT-5, on both VUE-TR-V2 and VUE-STG, while achieving competitive results with popular open-source models with similar scale on video QA benchmarks. The latest Vidi2.5 offers significantly stronger STG capability and slightly better TR and Video QA performance over Vidi2. This update also introduces a Vidi2.5-Think model to handle plot understanding with complex plot reasoning. To comprehensively evaluate the performance of plot understanding, we propose VUE-PLOT benchmark with two tracks, Character and Reasoning. Notably, Vidi2.5-Think outperforms Gemini 3 Pro Preview on fine-grained character understanding with comparable performance on complex plot reasoning. Furthermore, we demonstrate the effectiveness of Vidi2.5 on a challenging real-world application, video editing planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。