arXiv:2503.15005cs.CV2025-03CVPR被引 16

提出跨模态通用场景图,统一描述图像、文本等多源信息中的场景语义。

Universal Scene Graph Generation

  • 设计模块化架构,通过对象关联器缓解跨模态对齐差异。
  • 引入文本主导的对比学习,提升在不同领域下的生成性能。
  • 支持任意模态组合输入,适合多模态场景理解任务。

场景图(SG)表示能简洁高效地描述场景语义,推动了该领域的持续研究。现实世界中多种模态(如图像、文本、视频、3D数据)常共存,各自具有独特特征。然而,当前的场景图研究大多局限于单一模态建模,难以充分利用不同模态间的互补优势来刻画整体场景语义。为此,我们提出通用场景图(Universal SG, USG),一种能够从任意模态组合输入中全面表征场景语义的新范式,涵盖模态无关与模态特定的场景。进一步,我们设计了专用的USG解析器USG-Par,有效解决跨模态对象对齐和域外挑战两大瓶颈。USG-Par采用模块化架构实现端到端生成,其中引入对象关联器以缓解跨模态差距;同时提出以文本为中心的场景对比学习机制,通过将多模态对象与关系与文本场景图对齐,缓解领域不平衡问题。大量实验表明,相比独立的场景图,USG在表达场景语义方面更具优势,且所提的USG-Par在效率与性能上均表现更优。

原文摘要 · Abstract (English)

Scene graph (SG) representations can neatly and efficiently describe scene semantics, which has driven sustained intensive research in SG generation. In the real world, multiple modalities often coexist, with different types, such as images, text, video, and 3D data, expressing distinct characteristics. Unfortunately, current SG research is largely confined to single-modality scene modeling, preventing the full utilization of the complementary strengths of different modality SG representations in depicting holistic scene semantics. To this end, we introduce Universal SG (USG), a novel representation capable of fully characterizing comprehensive semantic scenes from any given combination of modality inputs, encompassing modality-invariant and modality-specific scenes. Further, we tailor a niche-targeting USG parser, USG-Par, which effectively addresses two key bottlenecks of cross-modal object alignment and out-of-domain challenges. We design the USG-Par with modular architecture for end-to-end USG generation, in which we devise an object associator to relieve the modality gap for cross-modal object alignment. Further, we propose a text-centric scene contrasting learning mechanism to mitigate domain imbalances by aligning multimodal objects and relations with textual SGs. Through extensive experiments, we demonstrate that USG offers a stronger capability for expressing scene semantics than standalone SGs, and also that our USG-Par achieves higher efficacy and performance.

场景图多模态跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。