arXiv:2509.19047cs.RO2025-09被引 14

用手持设备采集力觉数据,提升精密操作的机器人学习效果

ManipForce: Force-Guided Policy Learning with Frequency-Aware Representation for Contact-Rich Manipulation

  • 融合视觉与高频力觉信号的多模态Transformer架构
  • 六类真实任务平均成功率83%,显著优于纯视觉方法
  • 适合需要稳定接触力的精密装配场景

高接触频率的操作任务(如精密装配)需要精确控制交互力,但现有模仿学习方法主要依赖视觉示范。我们提出ManipForce,一种手持系统,可在自然人类示范过程中同步采集高频率力矩(F/T)与RGB数据。基于这些示范,我们引入频率感知多模态变换器(FMT),通过频率与模态感知嵌入编码异步的RGB和F/T信号,并在变换器扩散策略中使用双向交叉注意力融合。在六类真实世界高接触操作任务(如齿轮装配、盒子翻转、电池插入)上进行大量实验,基于ManipForce示范训练的FMT模型平均成功率达83%,显著优于仅使用RGB的基线方法。消融实验与采样频率分析进一步证实,引入高频F/T数据及跨模态融合能有效提升策略性能,尤其在高精度与稳定接触要求的任务中。

原文摘要 · Abstract (English)

Contact-rich manipulation tasks such as precision assembly require precise control of interaction forces, yet existing imitation learning methods rely mainly on vision-only demonstrations. We propose ManipForce, a handheld system designed to capture high-frequency force-torque (F/T) and RGB data during natural human demonstrations for contact-rich manipulation. Building on these demonstrations, we introduce the Frequency-Aware Multimodal Transformer (FMT). FMT encodes asynchronous RGB and F/T signals using frequency- and modality-aware embeddings and fuses them via bi-directional cross-attention within a transformer diffusion policy. Through extensive experiments on six real-world contact-rich manipulation tasks - such as gear assembly, box flipping, and battery insertion - FMT trained on ManipForce demonstrations achieves robust performance with an average success rate of 83% across all tasks, substantially outperforming RGB-only baselines. Ablation and sampling-frequency analyses further confirm that incorporating high-frequency F/T data and cross-modal integration improves policy performance, especially in tasks demanding high precision and stable contact.

力觉反馈多模态学习机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。