让视觉语言模型学会定位具体物体,提升时空理解能力
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
- 通过实例级对比学习,将文本描述与特定时空区域对齐
- 在200万图像、5万视频的数据集上,实例检索性能显著提升
- 适合需要精确定位和细粒度理解的应用场景
现有视觉语言预训练范式擅长全局场景理解,但在实例级推理上表现不佳,因仅依赖全局监督。我们提出InstAP,一种实例感知预训练框架,通过联合优化全局视觉-文本对齐与细粒度实例级对比对齐,将文本提及精准锚定到具体时空区域。为此,我们构建了大型数据集InstVL,包含200万张图像和5万段视频,具备双粒度标注:整体场景描述和密集的实例级接地描述。在InstVL基准测试中,InstAP在实例级检索任务上显著优于现有VLP模型,并超越在同一数据集上训练的强基线模型,凸显实例感知目标的优势。此外,以实例为中心的预训练也提升了全局理解能力:InstAP在多个视频基准(如MSR-VTT和DiDeMo)上实现具有竞争力的零样本性能。定性可视化显示,InstAP能准确将文本提及定位到对应实例,而全局模型则表现出更弥散的场景级注意力。
原文摘要 · Abstract (English)
Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only supervision. We introduce InstAP, an Instance-Aware Pre-training framework that jointly optimizes global vision-text alignment and fine-grained, instance-level contrastive alignment by grounding textual mentions to specific spatial-temporal regions. To support this, we present InstVL, a large-scale dataset (2 million images, 50,000 videos) with dual-granularity annotations: holistic scene captions and dense, grounded instance descriptions. On the InstVL benchmark, InstAP substantially outperforms existing VLP models on instance-level retrieval, and also surpasses a strong VLP baseline trained on the exact same data corpus, isolating the benefit of our instance-aware objective. Moreover, instance-centric pre-training improves global understanding: InstAP achieves competitive zero-shot performance on multiple video benchmarks, including MSR-VTT and DiDeMo. Qualitative visualizations further show that InstAP localizes textual mentions to the correct instances, while global-only models exhibit more diffuse, scene-level attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。