arXiv:2605.13386cs.LGstat.ML2026-05

交叉注意力生成模型本质是随时间变化的核平滑过程。

Support-Conditioned Flow Matching Is Kernel Smoothing

论文配图:Support-Conditioned Flow Matching Is Kernel Smoothing
图 1 · 摘自论文原文
  • 用核平滑理论解释交叉注意力的运动轨迹,揭示其随时间从广义平均到近邻加权的演化机制。
  • 在高维、几何不匹配或样本不足时,模型性能下降,与理论预测的三种失效情形完全一致。
  • 实验证明IP-Adapter等方法实际实现了近似核平滑,为设计更优条件生成模型提供理论依据。

生成模型常通过交叉注意力对少量示例进行条件化。在高斯最优传输路径下,我们证明有限支持集诱导的精确速度场是一个Nadaraya-Watson核平滑器,其带宽随流时间减小:早期为广义平均,后期趋近于最近邻。单个高斯核注意力头可精确计算该速度场,将交叉注意力与经典核理论直接关联。理论预测三种失效情形:高维下的最近邻坍缩、各向同性核与数据几何不匹配、非参数估计所需支持不足。在高斯混合、球面壳及DINOv2 ImageNet特征上的实验表明,学习到的条件化在这些情形下表现提升,且IP-Adapter的交叉注意力在实践中实现了近似NW平滑。

原文摘要 · Abstract (English)

Generative models are often conditioned on a small set of examples via cross-attention. Under the Gaussian optimal-transport path, we show that the exact velocity field induced by a finite support set is a Nadaraya--Watson kernel smoother whose bandwidth decreases with flow time, from broad averaging at early steps to nearest-neighbor at late steps. A single Gaussian-kernel attention head exactly computes this field, connecting cross-attention conditioning to classical kernel theory. The theory predicts three failure regimes: nearest-neighbor collapse of the kernel at high dimension, mismatch between the isotropic kernel and the data geometry, and insufficient support for nonparametric estimation. Experiments on Gaussian mixtures, spherical shells, and DINOv2 ImageNet features confirm that learned conditioning improves in precisely these regimes, and that IP-Adapter's cross-attention implements approximate NW smoothing in practice.

生成模型交叉注意力核平滑理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。