提出混合力-位控制策略,提升高精度接触操作的鲁棒性与安全性。
Learning Hybrid-Control Policies for High-Precision In-Contact Manipulation Under Uncertainty

- 动态选择每个维度使用力控或位控,实现更灵活的接触控制。
- 在定位误差下成功率提高10%,插销破损减少5倍,平均受力降低30%。
- 适合高不确定性环境下的精密装配任务,尤其适合机器人抓取与插入场景。
基于强化学习的控制策略在诸多操作任务中已证明优于解析方法。传统方法学习神经控制策略,直接从状态信息预测末端执行器位姿变化。然而,在需要施加力约束的精细操作(如插入脆弱连接器)中,位姿控制难以显式调控力,依赖精心调参的底层控制器避免破坏性动作。本文提出混合位置-力控制策略,学习在各控制维度上动态选择使用力控或位控。为提升学习效率,引入针对接触处理的模式感知训练(MATCH),使策略动作概率显式匹配混合控制中的模式选择行为。我们在极端定位不确定性下的易碎插销任务中验证了该方法的有效性。结果表明,MATCH显著优于仅位姿控制策略:成功率最高提升10%,插销断裂次数减少5倍;在常见状态估计误差下表现更优。此外,尽管在更大更复杂的动作空间中学习,其数据效率仍与位姿控制相当。在超过1600次模拟到真实实验中,面对高噪声环境(33%成功率对比68%),MATCH成功率达双倍;在实验室条件下,其施加的平均力比可变阻抗策略低约30%。
原文摘要 · Abstract (English)
Reinforcement learning-based control policies have been frequently demonstrated to be more effective than analytical techniques for many manipulation tasks. Commonly, these methods learn neural control policies that predict end-effector pose changes directly from observed state information. For tasks like inserting delicate connectors which induce force constraints, pose-based policies have limited explicit control over force and rely on carefully tuned low-level controllers to avoid executing damaging actions. In this work, we present hybrid position-force control policies that learn to dynamically select when to use force or position control in each control dimension. To improve learning efficiency of these policies, we introduce Mode-Aware Training for Contact Handling (MATCH) which adjusts policy action probabilities to explicitly mirror the mode selection behavior in hybrid control. We validate MATCH's learned policy effectiveness using fragile peg-in-hole tasks under extreme localization uncertainty. We find MATCH substantially outperforms pose-control policies -- solving these tasks with up to 10% higher success rates and 5x fewer peg breaks than pose-only policies under common types of state estimation error. MATCH also demonstrates data efficiency equal to pose-control policies, despite learning in a larger and more complex action space. In over 1600 sim-to-real experiments, we find MATCH succeeds twice as often as pose policies in high noise settings (33% vs.~68%) and applies ~30% less force on average compared to variable impedance policies on a Franka FR3 in laboratory conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。