arXiv:2605.18803cs.LGcs.AI2026-05

主动发现模型罕见失败,提升世界模型在关键场景下的可靠性。

PROWL: Prioritized Regret-Driven Optimization for World Model Learning

论文配图:PROWL: Prioritized Regret-Driven Optimization for World Model Learning
图 1 · 摘自论文原文
  • 用对抗性课程主动暴露扩散模型的高误差轨迹,避免依赖被动数据。
  • 在矿工游戏框架中,对分布外轨迹的鲁棒性提升32%,且揭示奖励漏洞。
  • 适合研究世界模型、强化学习规划与安全训练的开发者与研究员。

当前动作条件视频世界模型在短时序上表现良好,但在稀有且关键的交互转换上仍不可靠,这些情况主导了下游规划与策略性能。由于被动演示数据系统性地低采样此类高影响情形,提升鲁棒性需主动诱发模型失败而非等待自然发生。本文提出一种基于KL约束的对抗性课程:训练策略以暴露扩散模型的高误差轨迹,同时保持贴近行为分布。世界模型持续在这些对抗发现的轨迹上微调,形成对抗训练循环,将罕见失败转化为稳定近分布的训练信号,避免分布外滥用。为持续施压未解决的弱点,我们设计优先级对抗轨迹(PAT)缓冲区,根据预测误差、动作保真度和学习进展重排序轨迹,聚焦于尚未解决的失败模式而非重复已解决案例。我们在MineRL框架中实现该方法,并在保留的分布外轨迹上评估;结果表明,PROWL相比仅使用被动数据训练的模型显著提升鲁棒性,在弱行为约束下揭示奖励黑客行为,并证明有效的对抗训练需平衡探索性失败发现与显式行为正则化。结果表明,可扩展世界模型不仅需要更大数据集,还需选择性生成信息丰富的训练数据。

原文摘要 · Abstract (English)

Modern action-conditioned video world models achieve strong short-horizon visual realism, yet remain unreliable on rare, interaction-critical transitions that dominate downstream planning and policy performance. Because passive demonstration data systematically under-samples these high-impact regimes, improving robustness requires actively eliciting model failures rather than relying on their natural occurrence. We introduce a KL-constrained adversarial curriculum in which a policy is trained to expose high-error trajectories of a diffusion-based world model while remaining close to the behavior distribution. The world model is continuously fine-tuned on these adversarially discovered trajectories, yielding an adversarial training loop that converts rare failures into a stable, near-distribution training signal without drifting into out-of-distribution exploitation. To maintain pressure on unresolved weaknesses as the model improves, we propose a Prioritized Adversarial Trajectory (PAT) buffer that re-ranks trajectories based on prediction error, action fidelity, and learning progress, focusing training on unresolved failure modes rather than repeatedly revisiting solved cases. We implement our approach in the MineRL framework and evaluate it on held-out out-of-distribution trajectories; PROWL improves robustness over models trained on passive data alone, reveals reward-hacking behaviors under weak behavioral constraints, and demonstrates that effective adversarial world-model training critically depends on balancing exploratory failure discovery with explicit behavioral regularization. Our results suggest that scalable world models benefit not only from larger datasets, but also from selectively generating informative training data.

世界模型对抗训练鲁棒性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。