用合成数据提升机器人摆放物体的通用能力,成功率提升超30%。
GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
- 通过多模态大模型解析人类指令与视觉输入,生成结构化摆放计划
- 合成数据扩展使真实场景成功率提升30.04个百分点,兼具合理性与可行性
- 适合需要智能家居或服务机器人部署的研究者与工程师
机器人需在日常家务中协助人类整理物品,核心挑战在于物体摆放任务,需同时考虑语义偏好(如常识性物体关系)与几何可行性(如避免碰撞)。本文提出GOPLA,一种分层框架,通过增强的人类示范数据学习通用摆放策略。多模态大语言模型将人类指令和视觉输入转化为包含成对物体关系的结构化计划;空间映射器结合几何常识,将计划转为3D可操作地图;扩散式规划器基于测试时成本,生成受多计划分布和碰撞规避约束的摆放位姿。为解决数据稀缺问题,引入可扩展的合成数据生成管道,将人类示范拓展为多样化训练数据。大量实验表明,该方法在定位精度与物理合理性上相较第二名提升30.04个百分点,展现出在广泛真实机器人摆放场景中的强泛化能力。
原文摘要 · Abstract (English)
Robots are expected to serve as intelligent assistants, helping humans with everyday household organization. A central challenge in this setting is the task of object placement, which requires reasoning about both semantic preferences (e.g., common-sense object relations) and geometric feasibility (e.g., collision avoidance). We present GOPLA, a hierarchical framework that learns generalizable object placement from augmented human demonstrations. A multi-modal large language model translates human instructions and visual inputs into structured plans that specify pairwise object relationships. These plans are then converted into 3D affordance maps with geometric common sense by a spatial mapper, while a diffusion-based planner generates placement poses guided by test-time costs, considering multi-plan distributions and collision avoidance. To overcome data scarcity, we introduce a scalable pipeline that expands human placement demonstrations into diverse synthetic training data. Extensive experiments show that our approach improves placement success rates by 30.04 percentage points over the runner-up, evaluated on positioning accuracy and physical plausibility, demonstrating strong generalization across a wide range of real-world robotic placement scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。