用低成本机器人实现助人任务,无需编程就能学会
Improving the performance of AI-powered Affordable Robotics for Assistive Tasks
- 通过模仿视频学习动作,无需手动标注或编程
- 实测任务准确率超90%,比基线高40%以上
- 模型缩小5倍仍保持75%准确率,适合资源受限场景
到2050年,全球对辅助护理的需求预计达35亿人,远超人力供给。现有机器人成本高且需专业技术,难以普及。本文提出一种低成本机械臂,可完成喂食、清理洒物、取药等助人任务。系统采用演示视频的模仿学习,无需任务特定编程或人工标注。由六伺服电机、双摄像头与3D打印夹具构成。通过遥控操作收集了5万帧视频数据,涵盖三项任务。提出新型分段动作块变换器(PACT)捕捉时序依赖并分割运动动态,结合时间集成(TE)方法优化轨迹以提升精度与平滑性。在五种模型规模和四种架构上评估,经十小时真实环境测试,系统任务准确率超过90%,较基线最高提升40%。PACT实现5倍模型压缩,仍维持75%准确率。显著性分析显示系统依赖关键视觉线索,相位标记梯度在关键轨迹时刻达到峰值,表明具备有效时序推理能力。未来将探索双手协同与移动能力,拓展辅助功能。
原文摘要 · Abstract (English)
By 2050, the global demand for assistive care is expected to reach 3.5 billion people, far outpacing the availability of human caregivers. Existing robotic solutions remain expensive and require technical expertise, limiting accessibility. This work introduces a low-cost robotic arm for assistive tasks such as feeding, cleaning spills, and fetching medicine. The system uses imitation learning from demonstration videos, requiring no task-specific programming or manual labeling. The robot consists of six servo motors, dual cameras, and 3D-printed grippers. Data collection via teleoperation with a leader arm yielded 50,000 video frames across the three tasks. A novel Phased Action Chunking Transformer (PACT) captures temporal dependencies and segments motion dynamics, while a Temporal Ensemble (TE) method refines trajectories to improve accuracy and smoothness. Evaluated across five model sizes and four architectures, with ten hours of real-world testing, the system achieved over 90% task accuracy, up to 40% higher than baselines. PACT enabled a 5x model size reduction while maintaining 75% accuracy. Saliency analysis showed reliance on key visual cues, and phase token gradients peaked at critical trajectory moments, indicating effective temporal reasoning. Future work will explore bimanual manipulation and mobility for expanded assistive capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。