用生成模型从门洞视角推断隐藏房间的结构与语义,无需微调即可辅助机器人决策。
MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- 通过视觉语言模型引导补全、单目深度估计和语义分割,生成隐藏区域的3D点云假设。
- 在MatterDoor数据集上,生成的先验能准确预测目标物体出现概率与空间占据概率。
- 适用于无任务特定训练的机器人导航与抓取规划,尤其适合复杂室内环境。
自主机器人常仅能通过门洞部分观测房间,墙壁和场景结构遮挡了安全导航与目标导向动作所需的几何与语义信息。我们探究预训练生成视觉模型能否作为零样本离线先验,推断缺失的结构信息。这些先验需支持对未观测区域的时空语义查询,估计目标物体在隐藏区域的概率及该区域被占据的可能性。给定一个第一人称RGB图像和目标查询,我们的流程结合VLM引导的外推补全、单目深度估计和语义分割,采样隐藏房间的语义标注3D点云假设。我们提出了MatterDoor——基于Matterport3D构建的门洞遮挡室内场景基准数据集,并通过生成指标和模拟Stretch机器人的物体抓取任务评估生成先验。结果表明,无需针对具体任务微调,即可获得可用于规划的有用时空语义先验。
原文摘要 · Abstract (English)
Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant semantics needed for safe navigation and goal-directed action. We ask whether off-the-shelf pretrained generative vision models can derive this missing structure as zero-shot offline priors for robot reasoning. Such priors should support spatio-semantic queries over unobserved structure, estimating the target object likelihood in hidden regions and the probability that those regions are occupied. Given an egocentric RGB observation and target query, our pipeline uses VLM-guided outpainting, monocular depth estimation, and semantic segmentation to sample semantically labeled 3D point cloud hypotheses of the hidden room. We introduce MatterDoor, a Matterport3D-derived benchmark of doorway-occluded indoor scenes, and evaluate the resulting priors with generative metrics and simulated Stretch robot object-reaching tasks. Our results suggest that useful spatio-semantic priors for planning can be derived without problem-specific fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。