用AI模仿人类操作,提升工业精密抓取的自动化水平。
Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation
- 构建端到端多模态模仿学习框架,融合视觉、点云与力觉数据。
- 在数据中心线缆任务中达78%成功率,远超单摄像头基线的36%。
- 仅需100次示范即可训练,适合高稳定性工业场景部署。
精密操作仍是工业自动化的关键瓶颈,如电缆布线、连接器插入和精密切割等任务仍严重依赖人工。本文从传统模块化机器人流程迈向端到端多模态模仿学习框架,提出三项核心贡献:一套模拟数据中心电缆管理、汽车线束和齿轮箱装配的工业灵巧操作基准板(IDB);可扩展的模仿学习框架DAG-ROS;以及基于多模态扩散策略的AG-iDP3模型,融合RGB图像、点云、关节位置与腕部力矩数据。以数据中心线缆操作板为例,在48次试验中评估端到端AI策略性能,最佳配置为包含多视角RGB输入的多模态扩散策略(DP),使用R3M编码器处理图像,实现78%的抓取与插入综合成功率,显著优于单摄像头RGB DP基线的36%。所有配置仅需约100次遥操作示范即可完成训练。结果表明,正确学习的策略在鲁棒性、泛化性和部署效率上均优于传统视觉与控制方法,支持向高可用工业环境的可扩展机器人自动化转型。
原文摘要 · Abstract (English)
Dexterous manipulation remains a critical bottleneck in industrial automation; tasks such as cable routing, connector insertion, and precision assembly still rely heavily on manual labor despite decades of robotics research. This work presents a progression from classical, modular robotics pipelines toward an end-to-end multimodal imitation-learning framework for industrial dexterous manipulation. As a part of this work, we introduce three key contributions: a set of Industrial Dexterity Benchmark (IDB) boards aimed to mimic datacenter cable management, automotive cable harnesses, and gearbox assembly tasks; a scalable imitation learning framework (DAG-ROS); and a multimodal diffusion-based policy framework (AG-iDP3) that creates models fusing RGB images, point clouds, joint positions, and wrist-frame wrench data. Focusing on the datacenter cable manipulation board, we evaluate the performance of a task involving cleaning a single cable over variations of an end-to-end AI policy using 48 trials per configuration. The best performing configuration, a multimodal expansion Diffusion Policy (DP), includes a multi-view RGB image source passed through an R3M encoder and reaches a 78% grasp and insert combined task success rate. This performance marks a significant improvement over the 36% observed from the single-camera RGB DP baseline. Each of the tested configurations requires only approximately 100 teleoperated demonstrations per task phase. These results indicate that the correct learned policy can outperform classical vision and control robotic methods in robustness, generalization, and deployment efficiency, justifying a shift toward scalable robotic automation for high up-time industrial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。