arXiv:2601.08728cs.CV2026-01被引 2

通过迭代显著性估计提升场景图生成的公平性与空间理解能力

Salience-SGG: Enhancing Unbiased Scene Graph Generation with Iterative Salience Estimation

  • 引入迭代显著性解码器,聚焦具有显著空间结构的三元组
  • 在多个数据集上达到当前最优性能,显著提升稀有关系识别准确率
  • 适合关注模型公平性与空间推理的研究者和开发者

场景图生成(SGG)面临长尾分布问题,少数谓词类别占据主导,导致模型对罕见关系表现不佳。现有无偏SGG方法虽缓解了偏差,但常牺牲空间理解能力,过度依赖语义先验。本文提出Salience-SGG框架,引入迭代显著性解码器(ISD),强化具有显著空间结构的三元组。为此,我们设计了语义无关的显著性标签以指导ISD。在Visual Genome、Open Images V6和GQA-200上的实验表明,Salience-SGG实现领先性能,并在配对定位平均精度上显著优于现有无偏方法,证明其在空间理解上的提升。

原文摘要 · Abstract (English)

Scene Graph Generation (SGG) suffers from a long-tailed distribution, where a few predicate classes dominate while many others are underrepresented, leading to biased models that underperform on rare relations. Unbiased-SGG methods address this issue by implementing debiasing strategies, but often at the cost of spatial understanding, resulting in an over-reliance on semantic priors. We introduce Salience-SGG, a novel framework featuring an Iterative Salience Decoder (ISD) that emphasizes triplets with salient spatial structures. To support this, we propose semantic-agnostic salience labels guiding ISD. Evaluations on Visual Genome, Open Images V6, and GQA-200 show that Salience-SGG achieves state-of-the-art performance and improves existing Unbiased-SGG methods in their spatial understanding as demonstrated by the Pairwise Localization Average Precision

场景图生成无偏学习空间理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。