arXiv:2505.22159cs.ROcs.CV2025-05NeurIPS被引 115

让机器人通过力觉感知更精准完成插拔等精细操作

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

  • 将力觉信号作为核心模态,动态融合视觉语言信息
  • 在插拔任务中成功率提升至80%,较基线高23.2%
  • 适合需要触觉反馈的复杂抓取与装配场景

视觉-语言-动作(VLA)模型通过利用预训练的视觉和语言表征,推动了通用机器人操作的发展。然而,在涉及力控的接触密集型任务中,尤其在视觉遮挡或动态不确定性下,其表现受限。为此,我们提出ForceVLA,一种端到端操作框架,将外部力觉传感视为VLA系统中的第一类模态。ForceVLA引入FVLMoE——一种力觉感知的专家混合融合模块,在动作解码过程中动态整合预训练的视觉-语言嵌入与实时六轴力反馈。该机制实现跨模态专家的上下文感知路由,增强机器人对细微接触动态的适应能力。我们还构建了新数据集ForceVLA-Data,包含五个接触密集型操作任务中同步的视觉、本体感知与力矩信号。ForceVLA相较强基线pi_0模型平均任务成功率提升23.2%,在插拔任务中最高达80%。该方法凸显多模态融合对灵巧操作的重要性,并为物理智能机器人控制设立了新基准。代码与数据将在https://sites.google.com/view/forcevla2025发布。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncertainty. To address these limitations, we propose ForceVLA, a novel end-to-end manipulation framework that treats external force sensing as a first-class modality within VLA systems. ForceVLA introduces FVLMoE, a force-aware Mixture-of-Experts fusion module that dynamically integrates pretrained visual-language embeddings with real-time 6-axis force feedback during action decoding. This enables context-aware routing across modality-specific experts, enhancing the robot's ability to adapt to subtle contact dynamics. We also introduce \textbf{ForceVLA-Data}, a new dataset comprising synchronized vision, proprioception, and force-torque signals across five contact-rich manipulation tasks. ForceVLA improves average task success by 23.2% over strong pi_0-based baselines, achieving up to 80% success in tasks such as plug insertion. Our approach highlights the importance of multimodal integration for dexterous manipulation and sets a new benchmark for physically intelligent robotic control. Code and data will be released at https://sites.google.com/view/forcevla2025.

机器人操作力觉感知多模态融合灵巧操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。