用扩散模型+3D语义场景,让不同机器人跨平台通用抓取。
A Flexible Field-Based Policy Learning Framework for Diverse Robotic Systems and Sensors
- 基于扩散策略与D3Fields的3D语义表示,实现跨机器人操控泛化。
- 仅100次示范即达80%成功率,支持多传感器配置灵活切换。
- 适合做多机器人系统测试或远程操控的科研人员参考。
我们提出一种跨机器人视觉运动学习框架,融合基于扩散策略的控制与来自D3Fields的3D语义场景表示,实现操纵任务的类别级泛化。其模块化设计支持多种机器人摄像头配置,包括配备Microsoft Azure Kinect阵列的UR5机械臂和使用Intel RealSense传感器的双臂操作器,通过低延迟控制栈和直观遥操作实现。统一配置层可无缝切换不同硬件设置,便于数据采集、训练与评估。在抓取并举起积木任务中,该框架仅需100次示范即达到80%的成功率,展现出在不同平台和传感模态间的鲁棒技能迁移能力。这一设计为跨机器人泛化的大规模真实世界研究铺平了道路。
原文摘要 · Abstract (English)
We present a cross robot visuomotor learning framework that integrates diffusion policy based control with 3D semantic scene representations from D3Fields to enable category level generalization in manipulation. Its modular design supports diverse robot camera configurations including UR5 arms with Microsoft Azure Kinect arrays and bimanual manipulators with Intel RealSense sensors through a low latency control stack and intuitive teleoperation. A unified configuration layer enables seamless switching between setups for flexible data collection training and evaluation. In a grasp and lift block task the framework achieved an 80 percent success rate after only 100 demonstration episodes demonstrating robust skill transfer between platforms and sensing modalities. This design paves the way for scalable real world studies in cross robotic generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。