arXiv:2603.08342cs.RO2026-03被引 1

通过分阶段调度视觉与力控,实现高精度抓取与操作。

PhaForce: Phase-Scheduled Visual-Force Policy Learning with Slow Planning and Fast Correction for Contact-Rich Manipulation

  • 分阶段调度视觉与力控信号,慢规划快修正
  • 真实机器人任务平均成功率86%,较基线提升40个百分点
  • 适合复杂接触场景的机器人操作,尤其擅长微调与适应性

高接触密度的操作任务不仅需要以视觉为主导的任务语义,还需对力/扭矩瞬态进行闭环响应。然而,生成式视觉动作策略通常受限于推理延迟和动作分块,难以充分利用力控信号进行高频反馈。此外,现有力感知方法往往持续、无差别地注入力信息,缺乏对何时、何地、以何种强度施加力的明确调度机制。本文提出PhaForce,一种分阶段调度的视觉-力策略,通过统一的接触/阶段调度协调低频分块规划与高频残差修正。其包含:(i) 接触感知阶段预测器(CAP),估计接触概率与阶段信念;(ii) 慢速扩散规划器,采用双门控视觉-力融合与正交残差注入,在保留视觉语义的同时引入力信号;(iii) 快速修正器,基于阶段路由的残差在可解释的修正子空间中进行分块内微调。在多个真实机器人接触密集任务中,PhaForce平均成功率达86%(较基线+40个百分点),同时显著改善接触质量,有效调控交互力,并对分布外几何变化具有强鲁棒性。

原文摘要 · Abstract (English)

Contact-rich manipulation requires not only vision-dominant task semantics but also closed-loop reactions to force/torque (F/T) transients. Yet, generative visuomotor policies are typically constrained to low-frequency updates due to inference latency and action chunking, underutilizing F/T for control-rate feedback. Furthermore, existing force-aware methods often inject force continuously and indiscriminately, lacking an explicit mechanism to schedule when / how much / where to apply force across different task phases. We propose PhaForce, a phase-scheduled visual--force policy that coordinates low-rate chunk-level planning and high-rate residual correction via a unified contact/phase schedule. PhaForce comprises (i) a contact-aware phase predictor (CAP) that estimates contact probability and phase belief, (ii) a Slow diffusion planner that performs dual-gated visual--force fusion with orthogonal residual injection to preserve vision semantics while conditioning on force, and (iii) a Fast corrector that applies control-rate phase-routed residuals in interpretable corrective subspaces for within-chunk micro-adjustments. Across multiple real-robot contact-rich tasks, PhaForce achieves an average success rate of 86% (+40 pp over baselines), while also substantially improving contact quality by regulating interaction forces and exhibiting robust adaptability to OOD geometric shifts.

机器人操作力控策略分阶段调度视觉-力融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。