arXiv:2501.00195cs.LGcs.AI2025-01被引 2

发现世界模型的潜在误差可提升鲁棒性,提出新正则化方法增强训练稳定性。

Towards Unraveling and Improving Generalization in World Models

  • 将世界模型学习建模为随机动力系统,分析潜在表征误差的影响机制。
  • 零漂移误差下,适度误差能起隐式正则作用,提升模型鲁棒性。
  • 提出雅可比正则化方案,有效抑制非零漂移导致的误差累积问题。

世界模型在强化学习中表现出色,但其泛化与鲁棒性机制尚不清晰。本文将世界模型学习建模为随机微分方程,分析潜在表征误差对鲁棒性与泛化的影响,涵盖零漂移与非零漂移两种情形。理论与实验表明,在零漂移条件下,适度的潜在表示误差可作为隐式正则化,反而提升鲁棒性。针对非零漂移带来的误差累积问题,提出雅可比正则化策略,显著提升训练稳定性,加速收敛,并改善长时序预测精度。实验验证了该方法在多个视觉控制任务中的有效性。

原文摘要 · Abstract (English)

World models have recently emerged as a promising approach to reinforcement learning (RL), achieving state-of-the-art performance across a wide range of visual control tasks. This work aims to obtain a deep understanding of the robustness and generalization capabilities of world models. Thus motivated, we develop a stochastic differential equation formulation by treating the world model learning as a stochastic dynamical system, and characterize the impact of latent representation errors on robustness and generalization, for both cases with zero-drift representation errors and with non-zero-drift representation errors. Our somewhat surprising findings, based on both theoretic and experimental studies, reveal that for the case with zero drift, modest latent representation errors can in fact function as implicit regularization and hence result in improved robustness. We further propose a Jacobian regularization scheme to mitigate the compounding error propagation effects of non-zero drift, thereby enhancing training stability and robustness. Our experimental studies corroborate that this regularization approach not only stabilizes training but also accelerates convergence and improves accuracy of long-horizon prediction.

世界模型强化学习正则化鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。