让机器人在接触操作中感知并调控力,提升稳定性和精度。
ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation
- 用视觉语言模型构建力觉任务概念,实现力感知的端到端控制。
- 在5个接触任务中成功率比pi0高48.0%,比pi0.5高35.0%。
- 适合需要精细力控的机器人操作场景,如装配、擦拭等。
面向接触丰富的操作任务,现有方法多依赖位置控制,对交互力的显式感知与调节研究不足,制约了真实场景下的稳定性、精度与鲁棒性。本文提出ForceVLA2,一种端到端的视觉-语言-动作框架,赋予机器人混合力-位置控制能力与显式的力觉感知。ForceVLA2将力基提示引入视觉语言模型专家,构建跨阶段的力感知任务概念;并在动作专家中采用跨尺度混合专家(MoE)结构,自适应融合这些概念与实时交互力,实现闭环混合力-位置调控。为支持学习与评估,我们构建了ForceVLA2-Dataset,包含1000条轨迹,覆盖5个接触丰富任务(擦除、按压、组装等),含多视角图像、任务提示、本体状态及力信号。大量实验表明,ForceVLA2在5个任务上显著提升成功率与可靠性,相较pi0和pi0.5分别提升48.0%与35.0%,有效缓解机械臂过载与接触不稳定等常见失败模式,推动视觉-语言-动作系统在力感知交互物理智能方面的进展。
原文摘要 · Abstract (English)
Embodied intelligence for contact-rich manipulation has predominantly relied on position control, while explicit awareness and regulation of interaction forces remain under-explored, limiting stability, precision, and robustness in real-world tasks. We propose ForceVLA2, an end-to-end vision-language-action framework that equips robots with hybrid force-position control and explicit force awareness. ForceVLA2 introduces force-based prompts into the VLM expert to construct force-aware task concepts across stages, and employs a Cross-Scale Mixture-of-Experts (MoE) in the action expert to adaptively fuse these concepts with real-time interaction forces for closed-loop hybrid force-position regulation. To support learning and evaluation, we construct ForceVLA2-Dataset, containing 1,000 trajectories over 5 contact-rich tasks, including wiping, pressing, and assembling, with multi-view images, task prompts, proprioceptive state, and force signals. Extensive experiments show that ForceVLA2 substantially improves success rates and reliability in contact-rich manipulation, outperforming pi0 and pi0.5 by 48.0% and 35.0%, respectively, across the 5 tasks, and mitigating common failure modes such as arm overload and unstable contact, thereby actively advancing force-aware interactive physical intelligence in VLAs. The project page is available at https://sites.google.com/view/force-vla2/home.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。