arXiv:2608.06170cs.ROcs.CV2026-08

让机器人在复杂环境中自动划分功能区域,还能根据任务动态调整

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

论文配图:Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments
图 1 · 摘自论文原文
  • 用先验知识引导视觉与几何信息融合,解决无墙空间的区域划分难题
  • 在真实开放空间中实现最高精度的功能区分割,远超现有方法
  • 支持零样本任务切换,可按新目标重新组织空间结构

层次化3D场景图是自主移动平台进行高层空间推理的有力表示。然而,现有提取框架通常依赖纯局部视觉聚类或严格几何启发式规则(如墙分隔的房间),在开放式或任意结构环境中失效。本文提出Prior-SG,一种任务与先验驱动的框架,将场景图生成从根本上视为概率对齐问题。机器人探索时,持续将输入的RGB-D流整合为物理上一致的实例图,采用多尺度、开放词汇特征融合策略。系统通过最大后验(MAP)估计推断地图的高层次功能语义,由先验图引导——该图由大语言模型动态合成的环境结构预期和任务相关词汇构成。通过优化融合异构专家(视觉、几何、离散物体)与拓扑先验的马尔可夫随机场,系统解决局部感知歧义。我们在多种模拟住宅数据集及大型真实开放环境验证该方法。Prior-SG相比近期基线达到最先进的语义区域分割准确率,在缺乏物理墙的情况下仍能稳健划分远距离功能边界,并首次实现零样本本体灵活性,使机器人可根据给定高层任务完全重构其空间分区。

原文摘要 · Abstract (English)

Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments. We propose Prior-SG, a task- and prior-driven framework that casts scene graph generation fundamentally as a probabilistic alignment problem. As the robot explores, it continuously aggregates an incoming RGB-D sensor stream into a physically grounded Instance Graph utilizing a multi-scale, open-vocabulary feature fusion strategy. The system then infers the high-level functional semantics of this map through a Maximum A Posteriori (MAP) estimate, guided by a Prior Graph-a logical expectation of the environment's structure and task-relevant vocabulary synthesized dynamically by a Large Language Model. By optimizing a Markov Random Field that fuses heterogeneous experts (visual, geometric, and discrete objects) with these topological priors, the system resolves local perceptual ambiguities. We validate this approach across diverse simulated residential datasets and large, open-plan real-world environments. Prior-SG achieves state-of-the-art semantic region segmentation accuracy compared to recent baselines, robustly delineates distant functional boundaries in the absence of physical walls, and uniquely provides zero-shot ontological flexibility, enabling the robot to entirely restructure its spatial partitioning based on a given high-level task.

场景图空间理解机器人导航大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。