用超图建模物体间隐含关系,提升机器人规划准确率
GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning

- 用超图结构表示图像中物体与功能区域的语义关联
- 在ActPlan1K和ALFRED上成功率、路径正确率均显著提升
- 适合需要理解复杂场景逻辑的智能体规划任务
具身智能系统中的程序化规划依赖于物体属性所编码的隐式空间关系与深层语义结构。现有方法过度依赖视觉语言模型自身的推理能力,忽视了多模态输入中可挖掘的丰富结构化语义信息,导致难以有效理解复杂场景中的功能空间关系。为此,我们提出GaLa,一种用于多模态程序化规划的视觉语言框架。GaLa引入基于超图的表示:将图像中的物体实例作为节点,依据属性与功能语义聚合物体构建区域级超边,显式捕捉物体间的隐式语义关系及功能区域的层次组织。进一步设计了三视图超图编码器,通过对比学习在节点视图、区域视图及节点-区域关联视图间强制语义一致性,使超图语义更有效地注入下游视觉语言模型推理。在ActPlan1K和ALFRED基准上的大量实验表明,GaLa在执行成功率、最长共同子序列(LCS)和规划正确性方面显著优于现有方法。
原文摘要 · Abstract (English)
Implicit spatial relations and deep semantic structures encoded in object attributes are crucial for procedural planning in embodied AI systems. However, existing approaches often over rely on the reasoning capabilities of vision language models (VLMs) themselves, while overlooking the rich structured semantic information that can be mined from multimodal inputs. As a result, models struggle to effectively understand functional spatial relationships in complex scenes. To fully exploit implicit spatial relations and deep semantic structures in multimodal data, we propose GaLa, a vision language framework for multimodal procedural planning. GaLa introduces a hypergraph-based representation, where object instances in the image are modeled as nodes, and region-level hyperedges are constructed by aggregating objects according to their attributes and functional semantics. This design explicitly captures implicit semantic relations among objects as well as the hierarchical organization of functional regions. Furthermore, we design a TriView HyperGraph Encoder that enforces semantic consistency across the node view, area view, and node area association view via contrastive learning, enabling hypergraph semantics to be more effectively injected into downstream VLM reasoning. Extensive experiments on the ActPlan1K and ALFRED benchmarks demonstrate that GaLa significantly outperforms existing methods in terms of execution success rate, LCS, and planning correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。