分离条件理解与去噪,提升预测扩散模型采样一致性
Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
- 用独立流分别处理条件输入和目标去噪,避免信息纠缠
- 引入预训练预测器提取关键特征,生成更符合真实轨迹的样本
- 在机器人视频和科学时空预测上表现更优,适合需要高一致性的场景
扩散和基于流的模型在多模态生成任务中取得显著进展,并已应用于预测学习。然而,与鼓励样本多样性的生成任务不同,预测学习涉及多种随机性,要求采样结果与真实轨迹保持一致,而我们发现在扩散模型中存在采样不一致问题。本文认为,预测扩散模型采样不一致的关键瓶颈在于预测能力不足,源于条件理解与目标去噪在共享架构和联合训练下的耦合。为此,提出Foresight Diffusion(ForeDiff)框架,通过解耦条件理解与目标去噪,提升采样一致性。ForeDiff采用独立的确定性预测流处理条件输入,同时利用预训练预测器提取信息丰富的表示以指导生成。在机器人视频预测和科学时空预测任务上的大量实验表明,ForeDiff在预测准确率和采样一致性方面均优于强基线,为预测扩散模型提供了有前景的新方向。
原文摘要 · Abstract (English)
Diffusion and flow-based models have enabled significant progress in generation tasks across various modalities and have recently found applications in predictive learning. However, unlike typical generation tasks that encourage sample diversity, predictive learning entails different sources of stochasticity and requires sampling consistency aligned with the ground-truth trajectory, which is a limitation we empirically observe in diffusion models. We argue that a key bottleneck in learning sampling-consistent predictive diffusion models lies in suboptimal predictive ability, which we attribute to the entanglement of condition understanding and target denoising within shared architectures and co-training schemes. To address this, we propose Foresight Diffusion (ForeDiff), a framework for predictive diffusion models that improves sampling consistency by decoupling condition understanding from target denoising. ForeDiff incorporates a separate deterministic predictive stream to process conditioning inputs independently of the denoising stream, and further leverages a pretrained predictor to extract informative representations that guide generation. Extensive experiments on robot video prediction and scientific spatiotemporal forecasting show that ForeDiff improves both predictive accuracy and sampling consistency over strong baselines, offering a promising direction for predictive diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。