arXiv:2607.02269cs.CVcs.AI2026-07

为视觉语言模型视频定位设计专用领域评估基准,检验真实场景下的适应能力。

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

论文配图:AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
图 1 · 摘自论文原文
  • 构建五个专业领域视频数据集,含精细时空标注与训练子集
  • 15个顶尖模型在新领域上零样本和上下文学习均表现不佳
  • 揭示当前模型在复杂时空推理上的根本缺陷,适合研究领域适应的学者

视觉语言模型在时空视频定位任务中展现出巨大潜力,但现有评估仍局限于通用日常场景的零样本测试,难以反映真实应用中面对罕见视觉概念与复杂时空动态的问题。由于无法对无限数据分布进行全量预训练,模型在新领域的适应能力至关重要。为此,我们提出AnyGroundBench,一个面向领域自适应的视频定位基准,涵盖动物、工业、体育、手术和公共安全五个专业领域。该基准结合新采集的专家标注视频(如小鼠行为)与已有数据集,通过密集高保真时空标注统一整合,并提供专属训练子集以系统评估模型的领域适应能力。我们评估了15个前沿视觉语言模型,在实际计算约束下测试其零样本泛化与上下文学习能力。结果表明,现有模型在专业领域中无论是零样本还是基于上下文学习的适应均表现失败,暴露出严重的时空推理缺陷,亟需未来研究突破。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world applications in specialized fields, where models inevitably encounter rare visual concepts and complex spatio-temporal dynamics. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured videos such as expert-annotated mouse behaviors with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated training subsets to systematically measure domain adaptability. We extensively evaluate 15 state-of-the-art VLMs, assessing their zero-shot generalization and In-Context Learning (ICL) capabilities under practical computational constraints. Ultimately, our findings reveal that current models fail in both zero-shot and ICL-based adaptation when confronted with specialized domains, exposing critical flaws in spatio-temporal reasoning that future research must address.

视频定位领域适应多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。