用双向扩散模型自测生成误差,无需真实数据也能判断预测可靠性。
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors

- 训练双向扩散模型,正反推演后回溯差值作为误差自监督信号。
- 在天体物理和人脸视频上验证,误差排序相关性达0.91~0.98,校准精度超95%覆盖。
- 适合关注生成模型可信度、需无监督误差检测的科研与工程应用者。
自回归模型在长序列生成中会累积误差,但部署时缺乏真实标签进行评估。本文训练一个单一的条件隐空间扩散模型,通过方向标志实现系统正向或反向时间推进,并证明该双向性可提供免标注的测试时误差信号:向前推进i步再反向i步应返回起点,往返差异$ℓ_i$即为不可观测生成误差的自监督代理。在可压缩磁流体动力学(MHD)和自然人脸视频(CelebV-HQ)上验证,对保留的MHD轨迹,$ℓ_i$在固定深度下误差排序相关性达Spearman 0.91–0.98,且在轨迹内达0.69±0.16;仅用训练数据拟合的简单校准器即可将误差估计精度控制在1.14倍(68%)和1.29倍(95%)以内,接近无偏。该信号在分布外的Orszag-Tang涡旋中准确标记异常(AUROC 0.98;深度10时为1.0),且在80%覆盖率下将误差降低15%,是仅依赖深度基线的三倍。双向训练零成本代价,同时优于单向专用模型,反向生成还可作快速逆求解器。在LE-PDE-UQ湍流纳维-斯托克斯基准上,单一双向模型以十分之一训练成本达到十模型集成的1.3倍精度,且实现最优无训练像素级校准。往返一致性将可逆性转化为生成模型的实际可信信号。
原文摘要 · Abstract (English)
Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against. We train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a direction flag, and show that this bidirectionality supplies a measurement-free test-time error signal: rolling forward $i$ steps and then backward $i$ steps must return the model to its start, so the round-trip discrepancy $\mathcal{C}_i$ is a self-supervised proxy for the unobservable rollout error: no ensembles, no held-out data, no governing equations, for one extra rollout. We validate on compressible magnetohydrodynamics (MHD), an astrophysical turbulent radiative mixing layer, and natural face videos (CelebV-HQ). On held-out MHD trajectories, $\mathcal{C}_i$ ranks rollout error (Spearman $0.91$-$0.98$ at fixed depth; $0.69 \pm 0.16$ within trajectories), and a simple calibrator fit on training rollouts predicts its magnitude to within $1.14\times$ ($68\%$) and $1.29\times$ ($95\%$) with near-nominal coverage - one nat beyond a depth-only predictor, transferring to all six decoded physical fields. The same signal flags the out-of-distribution Orszag-Tang vortex (AUROC $0.98$; $1.0$ by depth $10$) exactly where sampling-dispersion baselines invert, and it cuts incurred error by $15\%$ at $80\%$ coverage - three times the depth-only baseline. Bidirectional training comes at negative cost, beating direction specialists in both directions, and the backward direction doubles as a fast inverse solver. On LE-PDE-UQ's turbulent Navier-Stokes benchmark, a single bidirectional model reaches accuracy within $1.3\times$ of their ten-model ensemble at a tenth of the training cost, with the best training-free pixel-level calibration. Round-trip consistency turns reversibility into a practical trust signal for generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。