arXiv:2605.11048cs.ROcs.AI2026-05被引 4

让机器人通过力觉与视觉协同完成复杂接触操作,提升成功率与泛化能力。

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

论文配图:ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching
图 1 · 摘自论文原文
  • 基于流匹配构建力觉感知的反应式控制框架,融合力信号与运动信息。
  • 在6个真实任务中成功率达37%提升,且在分布外场景下表现优异。
  • 适合需要高精度力控的机械臂操作任务,如装配、抓取等。

现有模仿学习方法使机器人能够自主与物理环境交互,但富含接触的操作任务仍面临挑战,因复杂的接触动力学需高精度力反馈与控制。尽管近期工作尝试将力/扭矩传感融入策略,如何构建一个简单有效、可在多模态观测下实现鲁棒泛化的框架仍是未解问题。本文提出ForceFlow,一种基于流匹配的力觉感知反应式框架。针对接触阶段策略设计,研究了力信号融合机制,采用不对称多模态融合架构,将力信号视为全局调控信号,并结合联合预测范式,增强策略对瞬时力和历史信息的理解,实现力与运动的深度耦合。针对任务级分层分解,将操作分为以视觉为主导的接近阶段(基于VLM的指针定位)和以触觉为主导的交互阶段(力驱动接触执行),并引入视觉到力觉(V2F)交接机制,显式解耦空间泛化与接触调控。六项真实世界接触密集型任务的实验结果表明,ForceFlow相比强基线ForceVLA成功率提升37%,同时成本显著更低。此外,ForceFlow表现出精确的力信号预测能力,在接触力自调节和零样本分布外(OOD)泛化方面表现卓越。

原文摘要 · Abstract (English)

Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulation tasks remain a significant challenge due to complex contact dynamics that demand high-precision force feedback and control. Although recent efforts have attempted to integrate force/torque sensing into policies, how to build a simple yet effective framework that achieves robust generalization under multimodal observations remains an open question. In this paper, we propose ForceFlow, a force-aware reactive framework built upon flow matching. For contact-stage policy design, we investigate force signal fusion mechanisms and adopt an asymmetric multimodal fusion architecture that treats force as a global regulatory signal, combined with a joint prediction paradigm that enhances the policy's understanding of instantaneous force and historical information, thereby achieving deep coupling between force and motion. For task-level hierarchical decomposition, we divide manipulation into a vision-dominant approach stage (VLM-based pointing for target localization) and a touch-dominant interaction stage (force-driven contact execution), with a Vision-to-Force (V2F) handover mechanism that explicitly decouples spatial generalization from contact regulation. Experimental results across six real-world contact-rich tasks demonstrate that ForceFlow achieves a 37% success rate improvement over the strong baseline ForceVLA while maintaining significantly lower cost. Moreover, ForceFlow exhibits accurate force signal prediction and demonstrates superior performance in contact force self-regulation and zero-shot out-of-distribution (OOD) generalization.

力觉控制模仿学习多模态融合机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。