揭示生成模型少步积分误差的空间传播机制,发现误差主要来自远处传递而非本地注入。
Spatial Transport of Integration Error in Generative ODEs

- 通过线性化动态传播逐步截断残差,追踪误差在图像中的空间传输路径。
- 实验显示90%以上误差源自其他区域,本地注入贡献不足10%。
- 模型内部速度场结构可预测误差分布,适合优化生成质量的算法研究者。
训练好的流模型或扩散模型通常仅用少量求解步骤,留下的积分误差在图像中分布不均。本文研究误差的注入位置及其如何传播至终点,提出一种带符号的源-传输会计方法,经256²分辨率下五种模型的一阶验证。扰动实验表明,学习到的动力学会广泛传播局部扰动:采样初期,超过90%的终点响应不再保留在原始位置。通过模型自身线性化动态传播的有符号单步截断残差,可重构终点误差的方向与区域结构(余弦相似度0.81-0.87)。一个区域的误差更多来自外部输入而非自身注入。结构破坏性零点实验进一步验证:随机化贡献符号使误差减半,重分配接收区域(保留内容、范数和符号)则彻底消除误差。误差注入位置可从模型本身读取。轨迹上速度或预测场的变化结构,在训练中自然形成,能预测最终各区域的误差差异(精细轨迹内图像相关系数0.57-0.70,仅用廉价求解时较弱)。该预测不完整,因终点误差不仅取决于注入量,还受符号、时机及通过学习动态的传输影响。对注入变化施加训练惩罚可降低少步误差,说明此结构可通过训练改变。
原文摘要 · Abstract (English)
A trained flow or diffusion model is usually run with only a handful of solver steps, and the integration error this leaves behind is unevenly distributed across the image. We ask where that error is injected and how it reaches the endpoint, and answer with a signed source-and-transport accounting of few-step integration error, tested to first order. A perturbation experiment on five models at 256^2 resolution shows the learned dynamics spread local disturbances widely: near the start of sampling, under 10% of the summed endpoint response remains at the source. Signed one-step truncation residuals, propagated through the model's own linearized dynamics, reconstruct much of the endpoint error's direction and regional structure (cosine 0.81-0.87), and a region's error owes more to what arrives from elsewhere than to its own injection. Structure-destroying nulls, with protocols frozen before evaluation, locate what carries the account: randomizing contribution signs halves it, and reassigning which region receives each contribution, with content, norms, and signs intact, destroys it entirely. Where the injections land is readable from the model itself. The variation of its velocity or prediction field along the trajectory, a structure that emerges during training, predicts the final per-region gap (within-image rho of 0.57-0.70 on fine trajectories, weaker from the cheap solve alone). The prediction is partial because endpoint error depends not only on injected magnitude but on its sign, timing, and transport through the learned dynamics. A training penalty on the injected variation lowers few-step error, so the structure is one a model can be trained to change.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。