arXiv:2601.18107cs.LGcs.HC2026-01

用可信合成数据提升离线强化学习性能,解决真实数据不足问题

Beyond Static Datasets: Robust Offline Policy Optimization via Vetted Synthetic Transitions

  • 通过双循环世界模型生成高质量合成轨迹,扩展训练数据
  • 在随机和次优数据上实现显著性能提升,最高达1.8倍改进
  • 多层不确定性过滤确保合成数据可靠性,适合工业机器人场景

离线强化学习(ORL)在工业机器人等安全关键领域具有巨大潜力,但静态数据集与策略间分布偏移导致需过度保守,限制性能提升。本文提出基于模型的MoReBRAC框架,通过不确定性感知的隐空间合成方法,克服此局限。该框架利用双循环世界模型生成高保真转移样本,扩充训练样本集。为保障合成数据可靠性,设计分层不确定性机制,融合变分自编码器(VAE)流形检测、模型敏感性分析与蒙特卡洛丢弃法,仅保留动态模型高置信区域内的转移样本。在D4RL Gym-MuJoCo基准测试中,尤其在“random”与“suboptimal”数据设置下表现优异,性能提升显著。研究还揭示了VAE作为几何锚点的作用,并讨论了从近最优数据学习时的分布权衡问题。

原文摘要 · Abstract (English)

Offline Reinforcement Learning (ORL) holds immense promise for safety-critical domains like industrial robotics, where real-time environmental interaction is often prohibitive. A primary obstacle in ORL remains the distributional shift between the static dataset and the learned policy, which typically mandates high degrees of conservatism that can restrain potential policy improvements. We present MoReBRAC, a model-based framework that addresses this limitation through Uncertainty-Aware latent synthesis. Instead of relying solely on the fixed data, MoReBRAC utilizes a dual-recurrent world model to synthesize high-fidelity transitions that augment the training manifold. To ensure the reliability of this synthetic data, we implement a hierarchical uncertainty pipeline integrating Variational Autoencoder (VAE) manifold detection, model sensitivity analysis, and Monte Carlo (MC) dropout. This multi-layered filtering process guarantees that only transitions residing within high-confidence regions of the learned dynamics are utilized. Our results on D4RL Gym-MuJoCo benchmarks reveal significant performance gains, particularly in ``random'' and ``suboptimal'' data regimes. We further provide insights into the role of the VAE as a geometric anchor and discuss the distributional trade-offs encountered when learning from near-optimal datasets.

离线强化学习合成数据不确定性建模机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。