让机器人在无全局视野下重排家具,挑战真实场景部署
Embodied Scene Rearrangement Planning

- 基于第一人称视角与顶层布局图进行长程规划
- 现有方法完成率不足,任务难度显著高于以往
- 适合研究具身智能与复杂任务规划的学者
本文提出具身场景重排规划(ESRP),要求具身智能体仅凭第一人称观测和顶层目标布局,在3D场景中重排家具以匹配目标配置。不同于以往任务,ESRP不提供全局状态信息,并引入物体间相互遮挡,更贴近真实机器人部署条件。这些因素使得将局部视角与全局目标对齐在长程规划中尤为困难。为此,我们构建了基于OmniGibson的ESRP-Bench基准,包含超过5,400个场景对和8,200个物体。定义了三级评估指标衡量重排质量,并提供了四种基线方法:分层任务与运动规划、视觉-语言模型驱动方法,以及两种基于学习的方法(监督学习与强化学习)。实验表明,当前方法难以高效完成任务,凸显ESRP作为具身智能在场景理解与长程任务规划中的前沿挑战。该工作为实现真实世界智能体部署奠定基础。
原文摘要 · Abstract (English)
This paper introduces Embodied Scene Rearrangement Planning (ESRP), a novel task requiring embodied agents to rearrange furniture in 3D scenes to match a target configuration using only egocentric observations and a top-down target layout. Unlike prior rearrangement tasks, ESRP precludes global state access and introduces mutual object occlusions, reflecting the practical constraints of real-world robotic deployment. These factors make aligning partial egocentric observations with the global target layout particularly challenging for long-horizon planning. To facilitate research, we present ESRP-Bench, a comprehensive benchmark built on OmniGibson featuring over 5,400 scene pairs and 8,200 objects. We define three multi-level metrics to evaluate rearrangement quality and provide four baselines: a hierarchical task-and-motion planning method, a vision-language-model-based method, and two learning-based approaches (IL and RL). Experimental results demonstrate that current methods struggle to complete the task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning. This work serves as a stepping stone toward deploying intelligent agents in real-world scenarios. Project page: https://pie-lab.cn/ESRP/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。