通过几何对称性干预,提升扩散模型的特征稳定性。
When Geometry Aligns: Dihedral Hidden-State Transformations in UNet, ViT, and DiT Architectures

- 用二面体群反射操作控制隐藏状态,统一分析多架构表现。
- 几何一致干预使特征更稳定,不一致则引发特定架构失效。
- 适用于研究图像生成中隐层动态与模型鲁棒性的学者。
当前扩散模型涵盖卷积式UNet与基于Transformer的结构(如扩散Transformer DiT),受视觉Transformer(ViT)启发,但其内部结构化几何扰动的影响仍不清楚。本文提出统一框架,通过二面体群的反射操作对中间隐藏状态进行可控干预,对比几何一致性与不一致变体。利用激活级诊断工具(自一致性偏移SCS、激活质量散度AMS、漂移量Drift)分析特征稳定性与几何漂移。结果显示,一致性变换显著提升稳定性,而不一致变换则引发可预测的、架构特异的失败。在Stable Diffusion 2.1 U-Net中,评估了七种干预模式,三个种子,结合图像级指标FID、KID、CLIP分数与LPIPS多样性验证。辅以ViT与受控DiT分析,确立几何一致性是空间结构化视觉与扩散模型中稳定隐藏状态干预的关键原则。
原文摘要 · Abstract (English)
Diffusion architectures now encompass convolutional UNets as well as transformer-based designs such as Diffusion Transformers (DiTs), inspired by Vision Transformers (ViTs), yet the effects of structured geometric perturbations within these architectures remain poorly understood. We study this question through a unified framework that applies reflection-based elements of the dihedral group to intermediate hidden states as controlled internal interventions, contrasting geometrically consistent and inconsistent variants. Using activation-level diagnostics, including Self-Consistency Shift (SCS), Activation Mass Scatter (AMS), and Drift, we analyze feature stability and geometric drift. We find that consistent transformations improve stability, while inconsistent ones induce predictable, architecture-specific failures. In the main Stable Diffusion 2.1 U-Net study, we evaluate seven intervention modes over three seeds and complement the internal diagnostics with image-level FID, KID, CLIP score, and LPIPS diversity. Taken together with supporting ViT and controlled DiT analyses, these results establish geometric consistency as a key principle for stable hidden-state interventions in spatially structured vision and diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。