让视觉关系生成更可信:用反事实验证避免语言偏见
CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation

- 通过反事实测试验证每个关系的视觉证据
- 在多个基准上提升未见关系泛化能力
- 适合关注可解释性和可靠性的视觉理解研究者
开放词汇场景图生成(SGG)旨在使用灵活且细粒度的关系短语描述视觉场景,突破固定谓词词汇表的限制。尽管最近的视觉-语言模型大幅扩展了SGG的语义覆盖范围,但也带来了可靠性问题:预测的关系可能受语言先验或物体共现驱动,而非真实视觉证据。本文提出一种基于反事实关系验证的证据强化型开放词汇SGG框架。不直接接受合理的候选关系,而是验证每个候选关系是否得到特定视觉、几何和上下文证据的支持。首先用视觉-语言提案器生成开放词汇关系候选,再将谓词短语分解为支持、接触、包含、深度和状态等软证据基元。关系条件证据编码器提取相关线索,反事实验证器则测试移除必要证据时关系得分是否下降,并在无关扰动下保持稳定。进一步引入矛盾感知谓词学习和图级偏好优化,以提升细粒度区分能力和全局图一致性。在常规、开放词汇及全景SGG基准上的实验表明,该方法在标准召回率、未见谓词泛化和反事实接地质量方面均持续提升。结果证明,从关系生成转向关系验证,能产生更可靠、可解释且证据奠基的场景图。
原文摘要 · Abstract (English)
Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate vocabulary. While recent vision-language models greatly expand the semantic coverage of SGG, they also introduce a critical reliability issue: predicted relations may be driven by language priors or object co-occurrence rather than grounded visual evidence. In this paper, we propose an evidence-rounded open-vocabulary SGG framework based on counterfactual relation verification. Instead of directly accepting plausible relation proposals, our method verifies whether each candidate relation is supported by relation-pecific visual, geometric, and contextual evidence. Specifically, we first generate open-vocabulary relation candidates with a vision-language proposer, then decompose predicate phrases into soft evidence bases such as support, contact, containment, depth and state. A relation-conditioned evidence encoder extracts predicate-relevant cues, while a counterfactual verifier tests whether the relation score decreases when necessary vidence is removed and remains stable under irrelevant perturbations. We further introduce contradiction-aware predicate learning and graph-level preference optimization to improve fine-grained discrimination and global graph consistency. Experiments on conventional, open-vocabulary, and panoptic SGG benchmarks show that our method consistently improves standard recall-based metrics, unseen predicate generalization, and counterfactual grounding quality. These results demonstrate that moving from relation generation to relation verification leads to more reliable, interpretable, and evidence-grounded scene graphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。