arXiv:2510.13358cs.ROcs.AI2025-10

通过对抗性微调提升机器人在故障下的控制鲁棒性

Adversarial Fine-tuning in Offline-to-Online Reinforcement Learning for Robust Robot Control

  • 在离线数据上训练后,注入动作扰动进行对抗微调
  • 相比纯离线方法,鲁棒性显著提升且收敛更快
  • 自适应课程策略可避免性能下降,适合实际部署

离线强化学习可实现无风险的高效策略训练,但基于静态数据集训练的策略在执行器故障等动作空间扰动下仍易失效。本文提出一种离线到在线的框架:先在干净数据上训练策略,再通过向执行动作注入扰动进行对抗性微调,促使策略学习补偿行为以增强鲁棒性。引入性能感知的课程策略,利用指数移动平均信号动态调整扰动概率,平衡鲁棒性与稳定性。连续控制行走任务实验表明,该方法在鲁棒性上持续优于纯离线基线,且收敛速度超过从零训练。匹配微调与评估条件可获得最强抗扰能力,自适应课程策略有效缓解线性课程导致的正常性能退化。结果表明,对抗性微调能实现不确定环境下的自适应与鲁棒控制,弥合了离线效率与在线适应性的差距。

原文摘要 · Abstract (English)

Offline reinforcement learning enables sample-efficient policy acquisition without risky online interaction, yet policies trained on static datasets remain brittle under action-space perturbations such as actuator faults. This study introduces an offline-to-online framework that trains policies on clean data and then performs adversarial fine-tuning, where perturbations are injected into executed actions to induce compensatory behavior and improve resilience. A performance-aware curriculum further adjusts the perturbation probability during training via an exponential-moving-average signal, balancing robustness and stability throughout the learning process. Experiments on continuous-control locomotion tasks demonstrate that the proposed method consistently improves robustness over offline-only baselines and converges faster than training from scratch. Matching the fine-tuning and evaluation conditions yields the strongest robustness to action-space perturbations, while the adaptive curriculum strategy mitigates the degradation of nominal performance observed with the linear curriculum strategy. Overall, the results show that adversarial fine-tuning enables adaptive and robust control under uncertain environments, bridging the gap between offline efficiency and online adaptability.

强化学习机器人控制鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。