arXiv:2601.20846cs.RO2026-01

用神经风格迁移生成真实数据,让机器人在少样本下学会切割未知材料。

End-to-end example-based sim-to-real RL policy transfer based on neural stylisation with application to robotic cutting

  • 将图像风格迁移思路迁移到时序数据,生成带物理真实感的仿真训练样本。
  • 仅需少量真实数据,任务完成时间更快,行为更稳定,优于基线方法。
  • 适合接触频繁、难获取奖励信号的复杂机器人任务,如未知材料切割。

尽管强化学习已在复杂不确定环境中的机器人控制中取得成功,但对大量数据(通常来自仿真环境)的依赖,限制了其在真实世界的部署,原因在于仿真与物理系统之间的领域差异,以及真实世界样本的稀缺性。本文提出一种基于神经风格迁移重构的新颖端到端模拟到现实的强化学习策略迁移方法,将图像处理中的风格迁移思想重新诠释,用于从无配对、无标签的真实世界数据集中合成新的训练数据。我们采用变分自编码器联合学习自监督特征表示,并生成弱配对的源-目标轨迹,以提升合成轨迹的物理真实性。我们在未知材料机器人切割这一案例研究中验证了该方法的有效性。相较于基线方法(包括我们的前期工作、CycleGAN 及基于条件变分自编码器的时间序列转换),本方法在极少真实数据条件下,实现了更短的任务完成时间和更高的行为稳定性。框架对几何与材料变化具有鲁棒性,证明了在缺乏真实世界奖励信息的高挑战性接触任务中策略适应的可行性。

原文摘要 · Abstract (English)

Whereas reinforcement learning has been applied with success to a range of robotic control problems in complex, uncertain environments, reliance on extensive data - typically sourced from simulation environments - limits real-world deployment due to the domain gap between simulated and physical systems, coupled with limited real-world sample availability. We propose a novel method for sim-to-real transfer of reinforcement learning policies, based on a reinterpretation of neural style transfer from image processing to synthesise novel training data from unpaired unlabelled real world datasets. We employ a variational autoencoder to jointly learn self-supervised feature representations for style transfer and generate weakly paired source-target trajectories to improve physical realism of synthesised trajectories. We demonstrate the application of our approach based on the case study of robot cutting of unknown materials. Compared to baseline methods, including our previous work, CycleGAN, and conditional variational autoencoder-based time series translation, our approach achieves improved task completion time and behavioural stability with minimal real-world data. Our framework demonstrates robustness to geometric and material variation, and highlights the feasibility of policy adaptation in challenging contact-rich tasks where real-world reward information is unavailable.

强化学习仿真到现实机器人切割风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。