arXiv:2502.04531cs.ROcs.AI2025-02中稿 · CoRL被引 18

用合成数据训练机器人在复杂场景中精准放置物体。

AnyPlace: Learning Generalized Object Placement for Robot Manipulation

  • 先用视觉语言模型定位大致区域,再聚焦局部预测精确姿态
  • 合成数据训练下成功率超基线,覆盖插入/堆叠/悬挂多种模式
  • 纯合成训练模型直接部署真实世界,适配多形态与高精度需求

机器人任务中的物体放置因物体几何形状和放置方式多样而极具挑战。为此,我们提出 AnyPlace,一种完全基于合成数据训练的两阶段方法,可预测多种真实任务中可行的放置姿态。核心思路是利用视觉语言模型(VLM)识别粗略放置位置,从而仅在相关区域进行局部姿态预测,使低层模型高效学习多样化放置策略。训练时,我们生成包含随机物体与不同放置方式(插入、堆叠、悬挂)的全合成数据集,并训练局部放置姿态预测模型。仿真评估表明,该方法在成功率、放置模式覆盖率和精度上均优于基线。真实世界实验显示,仅基于合成数据训练的模型可直接迁移到现实场景,在物体形态各异、放置方式多样且需高精度的条件下仍能成功执行放置任务。

原文摘要 · Abstract (English)

Object placement in robotic tasks is inherently challenging due to the diversity of object geometries and placement configurations. To address this, we propose AnyPlace, a two-stage method trained entirely on synthetic data, capable of predicting a wide range of feasible placement poses for real-world tasks. Our key insight is that by leveraging a Vision-Language Model (VLM) to identify rough placement locations, we focus only on the relevant regions for local placement, which enables us to train the low-level placement-pose-prediction model to capture diverse placements efficiently. For training, we generate a fully synthetic dataset of randomly generated objects in different placement configurations (insertion, stacking, hanging) and train local placement-prediction models. We conduct extensive evaluations in simulation, demonstrating that our method outperforms baselines in terms of success rate, coverage of possible placement modes, and precision. In real-world experiments, we show how our approach directly transfers models trained purely on synthetic data to the real world, where it successfully performs placements in scenarios where other models struggle -- such as with varying object geometries, diverse placement modes, and achieving high precision for fine placement. More at: https://any-place.github.io.

机器人操作放置预测合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。