VideoLoom统一建模视频时空理解,性能领先。
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
- 构建联合时空定位的视频大模型,引入精细化标注数据集
- 在多个基准上达到顶尖表现,如ReVOS上63.1的J&F分数
- 适合需要精准时空理解的视频分析任务,如智能监控与内容检索
本文提出VideoLoom,一种用于联合空间-时间理解的统一视频大语言模型。为提升细粒度空间与时间定位能力,我们构建了包含8.7k条视频的人类中心数据集LoomData-8.7k,其带有时间对齐和空间定位的描述。基于此,VideoLoom在多个空间与时间基准测试中取得领先或极具竞争力的表现(例如在ReVOS上达到63.1 J&F,Charades-STA上达到48.3 [email protected])。此外,我们提出LoomBench,一个包含时间、空间和组合式视频问答对的新基准,可从多角度全面评估视频大模型。这些贡献共同构建了一个通用且高效的联合时空视频理解体系,推动多模态智能新标准。
原文摘要 · Abstract (English)
This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we curate LoomData-8.7k, a human-centric video dataset with temporally grounded and spatially localized captions. With this, VideoLoom achieves state-of-the-art or highly competitive performance across a variety of spatial and temporal benchmarks (e.g., 63.1 J&F on ReVOS for referring video object segmentation, and 48.3 [email protected] on Charades-STA for temporal grounding). In addition, we introduce LoomBench, a novel benchmark consisting of temporal, spatial, and compositional video-question pairs, enabling a comprehensive evaluation of Video LLMs from diverse aspects. Collectively, these contributions offer a universal and effective suite for joint spatial-temporal video understanding, setting a new standard in multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。