arXiv:2502.03856cs.CV2025-02被引 1

让模型关注物体互动关系,提升场景图生成准确率

Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation

  • 通过区分互动与非互动物体,改进预训练目标生成
  • 细调阶段优先匹配互动对象对,提升关系识别精度
  • 引入一致性知识蒸馏,增强模型对真实互动的敏感性

当前开放词汇场景图生成(OVSGG)通过利用大规模预训练模型的知识,可识别未定义类别中的物体与关系。现有方法多采用两阶段流程:基于图像描述的弱监督预训练和全标注场景图的监督微调。然而,这些方法忽略物体间的互动关系,将所有物体同等对待,导致关系配对不匹配。为此,本文提出交互感知的OVSGG框架INOVA。在预训练阶段,使用交互感知的目标生成策略区分互动与非互动物体;在微调阶段,设计交互引导的查询选择机制,优先匹配互动物体对。此外,引入交互一致性的知识蒸馏,使互动对象对远离背景分布,增强鲁棒性。在VG和GQA两个基准上的实验表明,INOVA达到当前最优性能,验证了交互感知机制在真实应用中的潜力。

原文摘要 · Abstract (English)

Today's open vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Most existing methods adopt a two-stage pipeline: weakly supervised pre-training with image captions and supervised fine-tuning (SFT) on fully annotated scene graphs. Nonetheless, they omit explicit modeling of interacting objects and treat all objects equally, resulting in mismatched relation pairs. To this end, we propose an interaction-aware OVSGG framework INOVA. During pre-training, INOVA employs an interaction-aware target generation strategy to distinguish interacting objects from non-interacting ones. In SFT, INOVA devises an interaction-guided query selection tactic to prioritize interacting objects during bipartite graph matching. Besides, INOVA is equipped with an interaction-consistent knowledge distillation to enhance the robustness by pushing interacting object pairs away from the background. Extensive experiments on two benchmarks (VG and GQA) show that INOVA achieves state-of-the-art performance, demonstrating the potential of interaction-aware mechanisms for real-world applications.

场景图生成交互建模开放词汇视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。