世界模型幻觉可预测可预防,靠数据覆盖信号实现。
Hallucination in World Models is Predictable and Preventable

- 用数据覆盖率信号检测模型在状态-动作空间的盲区。
- 仅需50条真实轨迹即可适配未见环境,数据效率高。
- 适合做强化学习、机器人仿真等需要可靠模拟的研究者。
现代生成式世界模型能渲染越来越逼真的可控未来,但常出现幻觉:模拟过程视觉流畅却偏离真实动态。我们假设幻觉集中在状态-动作空间的低覆盖区域,轻量级数据驱动信号可同时检测并指导缓解。为此,我们构建了包含427小时、210项任务的MMBench2数据集,含真实动作、奖励和实时模拟器,并在此上训练了一个3.5亿参数的世界模型。我们识别出三种幻觉模式:感知型、动作边缘型和场景偏离型,分别对应流水线不同阶段,并开发了三种准确预测失败位置的信号。为填补训练时的覆盖空白,提出覆盖感知采样;在线则用幻觉预测作为好奇心奖励,进行针对性数据收集,实现仅用50条真实轨迹即完成预训练模型对全新环境的高效微调。结果表明,世界模型的幻觉本质上是数据覆盖问题,且检测信号本身也可用于缓解。
原文摘要 · Abstract (English)
Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that hallucination concentrates in low-coverage regions of the state-action space, where lightweight data-centric signals can both detect it and guide mitigation. To test this, we introduce MMBench2, a 427-hour, 210-task dataset for visual world modeling with ground-truth actions, rewards, and live simulators, and train a 350M-parameter world model on it. We identify three distinct hallucination modes: perceptual, action-marginalized, and scene-diverging -- each anchored to a different stage of the pipeline, and develop three signals that accurately predict where the model will fail. To close coverage gaps at training time, we develop a coverage-aware sampling technique; to close them online, our hallucination predictors serve as curiosity rewards for targeted data collection, yielding a data-efficient finetuning recipe that adapts the pretrained world model to entirely unseen environments with as few as 50 real environment trajectories. Overall, our findings reveal that hallucination in world models is inherently a data coverage issue, and that the same signals used to detect it can also be used for mitigation. An interactive web version of our paper is available at https://www.nicklashansen.com/mmbench2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。