提出球面自编码器,提升图像重建与生成质量。
Hyperspherical Autoencoder for High-Fidelity Image Reconstruction and Generation
- 分离语义方向与特征幅度,保留细节信息
- 球面流匹配训练扩散模型,实现高效收敛
- 在真实图像重建上达到25.2dB PSNR
近期研究探索使用DINO等预训练视觉基础模型构建生成自编码器,表现出优异的生成性能。然而,现有方法常因高频细节丢失而限制重建保真度。本文提出 extbf{超球面自编码器(HAE)},连接语义表征与像素级重建。核心洞察是:对比学习表示中语义信息主要体现为方向性,而严格匹配特征幅度会阻碍细粒度细节保留。为此,我们引入 extit{方向特征对齐}目标,在保持语义一致性的同时允许特征幅度灵活调整以保留细节,并设计 extit{分层卷积块嵌入}模块增强局部结构保持。此外,观察到自监督学习表征天然位于超球面上,我们采用 extit{黎曼流匹配}直接在该球形潜在流形上训练扩散变换器(DiT)。显著的是,该流形感知的DiT展现出极快收敛速度,在生成方面取得gFID extbf{1.96},重建方面达成rFID extbf{0.78}和PSNR extbf{25.2} dB,验证了流形感知方法的优势。
原文摘要 · Abstract (English)
Recent studies have explored using pretrained Vision Foundation Models (VFMs) such as DINO for generative autoencoders, showing strong generative performance. Unfortunately, existing approaches often suffer from limited reconstruction fidelity due to the loss of high-frequency details. In this work, we present the \textbf{\em Hyperspherical Autoencoder (HAE)}, a framework that bridges semantic representation and pixel-level reconstruction. Our key insight is that while semantic information in contrastive representations is primarily directional, enforcing strict magnitude matching hinders the preservation of fine-grained details. To address this, we introduce a {\em Directional Feature Alignment} objective that enforces semantic consistency while allowing flexible feature magnitudes for detail retention, alongside a {\em Hierarchical Convolutional Patch Embedding} module to enhance local structure preservation. Furthermore, observing that SSL-based representations intrinsically lie on a hypersphere, we employ {\em Riemannian Flow Matching} to train a Diffusion Transformer (DiT) directly on this spherical latent manifold. Notably, our manifold-aware DiT exhibits highly efficient convergence, achieving an exceptional gFID of \textbf{1.96} alongside a reconstruction rFID of \textbf{0.78} and a PSNR of \textbf{25.2} dB, validating the advantages of our manifold-aware approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。