提出双查询框架,融合两类图像理解方法提升场景图生成效果
Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability

- 设计双查询结构,同时利用检测器与查询两种推理方式
- 在三个数据集上均实现性能提升,最高准确率超基线3.2个点
- 适合需要高精度场景理解的视觉推理任务
场景图生成(SGG)方法可大致分为基于检测器和基于查询的两类,其内在推理机制导致预测行为存在差异,但这种差异尚未被系统分析。本文从检测器条件可达性视角出发,设计受控实验以探究预测差异。结果揭示了两类方法间存在明显互补性。受此启发,我们提出Dual-SGG方法,通过双查询设计融合两类推理机制,从而利用其互补的预测能力。在Visual Genome、Open Images v6和GQA-200数据集上的大量实验表明该方法有效,显著提升了场景图生成性能。
原文摘要 · Abstract (English)
Scene graph generation (SGG) approaches can be broadly classified into detector-based and query-based methods according to their underlying reasoning mechanisms. However, the discrepancy in their predictive behaviors, induced by these distinct mechanisms, has not been systematically analyzed. In this work, we design a controlled experimental setup to examine prediction discrepancies from the perspective of detector-conditioned reachability. The results suggest clear complementary clues. Motivated by this observation, we introduce a Dual-SGG method that consolidates both reasoning mechanisms via a dual-query design, thereby leveraging the complementary predictive behaviors of both detector-based and query-based methods. Extensive experiments on the Visual Genome, Open Images v6, and GQA-200 datasets demonstrate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。