让视频生成模型知道自己的不确定,避免幻觉错误。
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- 用严格评分规则训练模型,实现像素级不确定性估计。
- 在隐空间估算不确定性,提升稳定性并降低训练成本。
- 生成高分辨率热力图,直观标记不可信区域,适合机器人等安全场景。
可控视频生成模型虽能高保真合成视频,但常因幻觉导致未来帧与物理现实不符,影响机器人策略评估与规划。现有模型缺乏置信度评估能力,难以抑制幻觉。为此,本文提出C3方法,通过连续尺度校准的不确定性量化(UQ),在子块级别实现密集置信度估计,精确定位每帧中的不确定性。该方法创新性地采用严格合适的评分规则联合训练正确性与校准性;在隐空间中估算不确定性,避免像素空间带来的训练不稳与高昂成本;并将密集隐空间不确定性映射至可解释的像素级置信度,生成高分辨率热力图,直观展示不可信区域。在大规模机器人学习数据集Bridge和DROID上的大量实验及真实世界评估表明,该方法不仅在训练分布内提供校准的不确定性估计,还能有效检测分布外样本。
原文摘要 · Abstract (English)
Recent advances in generative video models have led to significant breakthroughs in high-fidelity video synthesis, specifically in controllable video generation where the generated video is conditioned on text and action inputs, e.g., in instruction-guided video editing and world modeling in robotics. Despite these exceptional capabilities, controllable video models often hallucinate - generating future video frames that are misaligned with physical reality - which raises serious concerns in many tasks such as robot policy evaluation and planning. However, state-of-the-art video models lack the ability to assess and express their confidence, impeding hallucination mitigation. To rigorously address this challenge, we propose C3, an uncertainty quantification (UQ) method for training continuous-scale calibrated controllable video models for dense confidence estimation at the subpatch level, precisely localizing the uncertainty in each generated video frame. Our UQ method introduces three core innovations to empower video models to estimate their uncertainty. First, our method develops a novel framework that trains video models for correctness and calibration via strictly proper scoring rules. Second, we estimate the video model's uncertainty in latent space, avoiding training instability and prohibitive training costs associated with pixel-space approaches. Third, we map the dense latent-space uncertainty to interpretable pixel-level uncertainty in the RGB space for intuitive visualization, providing high-resolution uncertainty heatmaps that identify untrustworthy regions. Through extensive experiments on large-scale robot learning datasets (Bridge and DROID) and real-world evaluations, we demonstrate that our method not only provides calibrated uncertainty estimates within the training distribution, but also enables effective out-of-distribution detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。