构建可感知状态的统一场景图,助力机器人理解家居环境并规划操作。
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
- 融合空间与功能关系,建模物体状态和可操作部件。
- 在基准测试中达71.6%准确率,较最佳基线提升11.4%。
- 适用于具身任务规划,支持零样本迁移至真实机器人。
家用移动操作机器人需同时完成导航与操作任务,这要求一种紧凑且语义丰富的场景表征,以捕捉物体位置、功能及其可操作部分。场景图是理想选择,但以往工作常将空间与功能关系分离,将场景视为静态快照,忽略对象状态与时间动态,且未关注任务相关的关键信息。为此,我们提出MomaGraph,一种集成空间-功能关系与部件级交互元素的统一场景表征。为推进该表征,我们构建了首个大规模、任务驱动的家居环境场景图数据集MomaGraph-Scenes,以及涵盖六项推理能力的系统评估基准MomaGraph-Bench。基于此,我们进一步开发了在MomaGraph-Scenes上通过强化学习训练的7B视觉语言模型MomaGraph-R1,能预测任务导向的场景图,并在Graph-then-Plan框架下实现零样本任务规划。大量实验表明,该模型在开源模型中达到领先水平,在基准上取得71.6%准确率(较最佳基线提升11.4%),并在公开基准上展现良好泛化能力,有效迁移到真实机器人实验。
原文摘要 · Abstract (English)
Mobile manipulators in households must both navigate and manipulate. This requires a compact, semantically rich scene representation that captures where objects are, how they function, and which parts are actionable. Scene graphs are a natural choice, yet prior work often separates spatial and functional relations, treats scenes as static snapshots without object states or temporal updates, and overlooks information most relevant for accomplishing the current task. To address these limitations, we introduce MomaGraph, a unified scene representation for embodied agents that integrates spatial-functional relationships and part-level interactive elements. However, advancing such a representation requires both suitable data and rigorous evaluation, which have been largely missing. We thus contribute MomaGraph-Scenes, the first large-scale dataset of richly annotated, task-driven scene graphs in household environments, along with MomaGraph-Bench, a systematic evaluation suite spanning six reasoning capabilities from high-level planning to fine-grained scene understanding. Built upon this foundation, we further develop MomaGraph-R1, a 7B vision-language model trained with reinforcement learning on MomaGraph-Scenes. MomaGraph-R1 predicts task-oriented scene graphs and serves as a zero-shot task planner under a Graph-then-Plan framework. Extensive experiments demonstrate that our model achieves state-of-the-art results among open-source models, reaching 71.6% accuracy on the benchmark (+11.4% over the best baseline), while generalizing across public benchmarks and transferring effectively to real-robot experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。