arXiv:2601.21633cs.CV2026-01

发现自编码器评估偏差,重建质量更关乎可控生成效果。

A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion

  • 提出用重建精度而非生成指标评估自编码器
  • 发现重建指标比gFID更准确预测可控性
  • 适合关注可控扩散模型的开发者与研究者

在潜在扩散模型中,自编码器(AE)通常需平衡重建保真度与生成友好潜空间(如低gFID)。近期基于ImageNet的AE研究显示,评估存在系统性偏向:重建指标被逐渐忽略,消融实验常选择gFID最优但重建退化的配置。我们理论分析表明,这种以gFID为主导的偏好在ImageNet生成任务中看似无害,但在扩展至可控扩散时风险凸显——自编码器可能引发条件漂移,限制条件对齐能力。我们发现,尤其在实例级重建指标上,其更能反映可控性。通过多维度条件漂移评估协议验证数个ImageNet规模AE,结果表明gFID仅弱预测条件保留,而重建导向指标显著更相关。ControlNet实验进一步证实,可控性与条件保留相关,而非gFID。本研究揭示了现有ImageNet中心评估体系与可扩展可控扩散需求间的差距,为更可靠的基准测试与模型选择提供指导。

原文摘要 · Abstract (English)

In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID). In recent ImageNet-scale AE studies, we observe a systematic bias toward generative metrics in handling this trade-off: reconstruction metrics are increasingly under-reported, and ablation-based AE selection often favors the best-gFID configuration even when reconstruction fidelity degrades. We theoretically analyze why this gFID-dominant preference can appear unproblematic for ImageNet generation, yet becomes risky when scaling to controllable diffusion: AEs can induce condition drift, which limits achievable condition alignment. Meanwhile, we find that reconstruction fidelity, especially instance-level measures, better indicates controllability. We empirically validate the impact of tilted autoencoder evaluation on controllability by studying several recent ImageNet AEs. Using a multi-dimensional condition-drift evaluation protocol reflecting controllable generation tasks, we find that gFID is only weakly predictive of condition preservation, whereas reconstruction-oriented metrics are substantially more aligned. ControlNet experiments further confirm that controllability tracks condition preservation rather than gFID. Overall, our results expose a gap between ImageNet-centric AE evaluation and the requirements of scalable controllable diffusion, offering practical guidance for more reliable benchmarking and model selection.

扩散模型可控生成自编码器条件漂移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。