arXiv:2505.06663cs.CV2025-05IJCAI被引 2

提出统一框架,让物体与关系相互增强,提升视频关系检测效果

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection

  • 用联合查询机制同时建模物体与关系,避免错误传播
  • 在VidVRD和VidOR上达到最新最好性能
  • 适合开放词汇视频理解任务研究者参考

开放词汇视频视觉关系检测旨在不限定预定义类别的情况下检测视频中的物体及其关系。现有方法依赖如CLIP等预训练视觉语言模型的丰富语义知识来识别新类别,通常采用级联流程先检测物体再分类关系,易导致误差传播,影响性能。本文提出一种基于查询的统一框架METOR,实现开放词汇场景下物体检测与关系分类的联合建模与相互增强。首先设计基于CLIP的上下文精炼编码模块,提取物体与关系的视觉上下文以优化文本特征与物体查询编码,提升对新类别的泛化能力;随后提出迭代增强模块,通过充分挖掘物体与关系间的相互依赖性,交替优化其表征,提升识别性能。在两个公开数据集VidVRD和VidOR上的大量实验表明,该框架取得当前最优表现。

原文摘要 · Abstract (English)

Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categories. Existing methods leverage the rich semantic knowledge of pre-trained vision-language models such as CLIP to identify novel categories. They typically adopt a cascaded pipeline to first detect objects and then classify relationships based on the detected objects, which may lead to error propagation and thus suboptimal performance. In this paper, we propose Mutual EnhancemenT of Objects and Relationships (METOR), a query-based unified framework to jointly model and mutually enhance object detection and relationship classification in open-vocabulary scenarios. Under this framework, we first design a CLIP-based contextual refinement encoding module that extracts visual contexts of objects and relationships to refine the encoding of text features and object queries, thus improving the generalization of encoding to novel categories. Then we propose an iterative enhancement module to alternatively enhance the representations of objects and relationships by fully exploiting their interdependence to improve recognition performance. Extensive experiments on two public datasets, VidVRD and VidOR, demonstrate that our framework achieves state-of-the-art performance.

视频关系检测开放词汇视觉语言模型联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。