arXiv:2503.09416cs.CV2025-03被引 2

用提示学习对齐语义空间,提升视频关系检测的泛化能力

OpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment

  • 通过区域描述生成文本表征,结合跨模态时空信息建模
  • 在VidVRD和VidOR数据集上优于现有方法,显著提升开放词汇检测性能
  • 适合研究视频理解、视觉关系检测与大模型应用的开发者

视频视觉关系检测(VidVRD)旨在识别视频中的物体及其关系,由于内容动态变化、标注成本高及关系分布长尾,任务极具挑战。尽管视觉语言模型(VLMs)有助于开放词汇检测,但常忽略视觉区域间的关系关联。且图像与视频间差异大,直接应用VLM检测视频关系困难。为此,我们提出OpenVidVRD框架,利用提示学习将VLM的知识迁移至视频关系检测。首先,基于视频区域自动生成区域描述,由VLM提取文本表征;其次,设计时空精炼模块,融合跨模态时空互补信息,生成对象级关系表征;最后,采用提示驱动的语义空间对齐策略,充分挖掘VLM的语义理解能力,增强模型泛化性。在VidVRD和VidOR公开数据集上的大量实验表明,所提方法优于现有方法。

原文摘要 · Abstract (English)

The video visual relation detection (VidVRD) task is to identify objects and their relationships in videos, which is challenging due to the dynamic content, high annotation costs, and long-tailed distribution of relations. Visual language models (VLMs) help explore open-vocabulary visual relation detection tasks, yet often overlook the connections between various visual regions and their relations. Moreover, using VLMs to directly identify visual relations in videos poses significant challenges because of the large disparity between images and videos. Therefore, we propose a novel open-vocabulary VidVRD framework, termed OpenVidVRD, which transfers VLMs' rich knowledge and powerful capabilities to improve VidVRD tasks through prompt learning. Specificall y, We use VLM to extract text representations from automatically generated region captions based on the video's regions. Next, we develop a spatiotemporal refiner module to derive object-level relationship representations in the video by integrating cross-modal spatiotemporal complementary information. Furthermore, a prompt-driven strategy to align semantic spaces is employed to harness the semantic understanding of VLMs, enhancing the overall generalization ability of OpenVidVRD. Extensive experiments conducted on the VidVRD and VidOR public datasets show that the proposed model outperforms existing methods.

视频理解关系检测提示学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。