arXiv:2410.22626cs.CV2024-10

通过联合图搜索理解复杂场景中的物体关系与空间布局。

Symbolic Graph Inference for Compound Scene Understanding

  • 构建场景与知识图谱联合推理,捕捉物体空间关系。
  • 在ADE20K数据集上优于现有方法,提升场景理解精度。
  • 适合需要细粒度场景语义解析的研究者与开发者。

场景理解是问答系统、机器人等多个领域的基础能力。与需显式学习相同场景不同组合的端到端方法不同,本方法通过分析场景中物体的构成及其排列方式来推断场景意义。提出一种新方法,在联合图搜索中同时利用场景图和知识图,捕获空间信息并整合通用领域知识。实验表明该方法在ADE20K数据集上可行,并优于当前主流场景理解方法。

原文摘要 · Abstract (English)

Scene understanding is a fundamental capability needed in many domains, ranging from question-answering to robotics. Unlike recent end-to-end approaches that must explicitly learn varying compositions of the same scene, our method reasons over their constituent objects and analyzes their arrangement to infer a scene's meaning. We propose a novel approach that reasons over a scene's scene- and knowledge-graph, capturing spatial information while being able to utilize general domain knowledge in a joint graph search. Empirically, we demonstrate the feasibility of our method on the ADE20K dataset and compare it to current scene understanding approaches.

场景理解图神经网络知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。