arXiv:2605.21688cs.ROcs.SY2026-05

用闭环反馈让仿真训练的策略直接控制微纤维形状,无需重训。

Closed-Loop Sim-to-Real Reinforcement Learning for Deformable Microfiber Shape Control

  • 在无摩擦简化仿真中训练强化学习策略,部署时靠视觉反馈实时修正接触误差。
  • 24种初始形态下平均点误差270±80μm,9个不同规格纤维均实现亚毫米级精度。
  • 适合做微尺度柔性物体操控的科研人员和工程师参考。

自主接触式微操纵因微尺度表面与界面相互作用难以精确建模,限制了传统基于模型的控制和仿真到现实的学习应用。本文提出一种用于表面微纤维形状调控的闭环仿真到现实强化学习方法。核心思想是在简化无摩擦仿真器中训练几何形状调控策略,并在实际部署时依赖实时视觉反馈,迭代修正未建模表面相互作用的影响。一个完全在仿真中训练的强化学习策略被直接应用于40 Hz运行的双机械臂微操纵系统,无需重新训练或领域自适应。以丝质微纤维为测试对象,该策略在24种不同初始配置下实现了270±80μm的平均点状形状误差。在涵盖三种直径(50、80、120μm)和三种操控长度(10、15、20mm)组合的九个样本上,同一策略无需重新训练或调参即达到亚毫米级最终形状误差。结果表明,只要任务相关的仿真-现实差异效应在闭环反馈中可观测且可纠正,仅在简化仿真中训练的策略即可实现可重复的真实世界微纤维形状调控。

原文摘要 · Abstract (English)

Autonomous contact-based micromanipulation is challenging because surface and interfacial interactions at the microscale are difficult to model accurately, limiting the use of conventional model-based control and sim-to-real learning. We present a closed-loop sim-to-real reinforcement learning (RL) approach for microfiber shape control on a surface. The central idea is to train geometric shape regulation in a simplified frictionless simulator and rely on real-time visual feedback during deployment to iteratively correct the observed effects of unmodeled surface interactions. An RL policy trained entirely in simulation is transferred directly to a physical dual-gripper micromanipulation system operating at 40 Hz, without retraining or domain adaptation. Using silk microfibers as a testbed, the policy achieves a mean point-wise shape error of 270 $\pm$ 80 $μ$m across twenty-four diverse initial configurations. Across nine specimens covering all combinations of three fiber diameters (50, 80, and 120 $μ$m) and three manipulated lengths (10 mm, 15mm, and 20 mm), the same policy achieves sub-millimeter final shape error without any retraining or retuning. These results show that a policy learned in a simplified simulator can achieve repeatable real-world microfiber shape regulation under surface contact, provided that the task-relevant effects of the sim-to-real mismatch remain observable and correctable within the closed feedback loop.

微操纵强化学习仿真到现实闭环控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。