arXiv:2607.14180cs.LGcs.AI2026-07

用人类偏好修复世界模型的幻觉问题,提升离线强化学习的可靠性。

RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences

论文配图:RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
图 1 · 摘自论文原文
  • 基于人类对想象轨迹的偏好,直接优化世界模型动态
  • 在多个环境上显著降低模型被利用的风险,减少灾难性遗忘
  • 适合希望低成本改进模型安全性的研究者与工程师

世界模型广泛用于离线强化学习以提升样本效率并生成超出固定数据集的经验。然而,当数据覆盖稀疏时,它们容易遭受模型利用问题。以往方法或依赖更多专家演示(成本高、不安全或不可得),或采用保守算法避开不确定区域(限制泛化能力)。本文提出直接通过人类对想象轨迹的偏好修复利用问题,利用人类对物理规律的直观判断识别明显错误的动力学幻觉。我们形式化为从人类反馈中学习动力学(DLHF),基于轨迹似然的Bradley-Terry偏好损失。但朴素的DLHF样本效率低,因此引入RENEW,利用认知不确定性聚焦微调最易被利用的区域。在Jumanji和经典控制环境中评估表明,尽管朴素DLHF需要大量偏好数据,RENEW显著提升样本效率,抑制灾难性遗忘,并减少预训练世界模型的利用现象。结果初步证明,人类偏好可直接监督世界模型动态,为解决离线模型基强化学习中的利用问题提供新思路。

原文摘要 · Abstract (English)

World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset. However, they are vulnerable to model exploitation where data coverage is thin. Prior work addresses this either by collecting more expert demonstrations, which is often expensive, unsafe, or unavailable, or by conservative algorithms that avoid uncertain regions, which limits generalization. We propose instead to repair exploitation directly using human preferences over imagined rollouts, leveraging the strong intuitive physics that allows humans to easily spot egregious dynamics hallucinations. We formalize this as Dynamics Learning from Human Feedback (DLHF), a Bradley-Terry preference loss over trajectory log-likelihoods under a learned dynamics model. Unfortunately, naive DLHF is sample inefficient, so we introduce RENEW, which uses epistemic uncertainty to focus finetuning where the model is most exploitable. We evaluate on several Jumanji and classic control environments and find that while naive DLHF requires an outsize preference budget, RENEW makes the framework practical by improving sample efficiency, limiting catastrophic forgetting, and reducing exploitation in pretrained world models. Taken together, our results provide initial evidence that preferences can supervise world model dynamics directly, offering a new approach to addressing exploitation in offline model-based RL.

世界模型强化学习偏好学习模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。