arXiv:2606.31290cs.LG2026-06被引 1

用局部正交分解构建可解释的低维潜在空间,实现高效超分辨率与精准不确定性量化。

Patch-PODiff-ViT: Structured Latent Diffusion with Patchwise POD for Super-Resolution and Uncertainty Quantification

论文配图:Patch-PODiff-ViT: Structured Latent Diffusion with Patchwise POD for Super-Resolution and Uncertainty Quantification
图 1 · 摘自论文原文
  • 以局部正交分解定义固定正交基,生成有序低维潜在表示
  • 参数更少、内存更低,且不确定性与真实集成结果高度吻合
  • 适合需要可解释不确定性的医学影像与气候数据超分辨任务

扩散模型能实现概率性超分辨率与条件生成,但像素空间方法计算成本高,而学习的潜在空间通常缺乏可解释的不确定性量化。我们提出Patch-PODiff-ViT,一种基于分块正交分解(POD)的结构化潜在扩散框架,其潜在空间由局部块上的固定线性正交基定义,而非由非线性自编码器学习。该方法生成低维、方差有序的潜在标记,保留空间结构,并支持在结构化低维潜在空间中通过视觉变换器高效扩散。由于解码器是固定、线性且正交的,潜在系数的不确定性可直接传播至物理空间预测方差,无需在像素空间进行蒙特卡洛估计即可实现解析传播。在海表温度、医学影像和自然图像上,该方法以更少参数和更低内存实现强重建效果,同时生成与经验集成高度匹配的校准空间不确定性。

原文摘要 · Abstract (English)

Diffusion models enable probabilistic super-resolution and conditional generation, but pixel-space methods are computationally expensive and learned latent spaces often lack interpretable uncertainty quantification. We introduce Patch-PODiff-ViT, a structured latent diffusion framework in which the latent space is defined by patchwise Proper Orthogonal Decomposition (POD), a fixed linear orthonormal basis over local patches, rather than learned by a nonlinear autoencoder. This yields low-dimensional, variance-ordered tokens that preserve spatial structure and enable efficient diffusion in a structured low-dimensional latent space with a Vision Transformer. Because the decoder is fixed, linear, and orthonormal, latent coefficient uncertainty can be propagated directly to physical-space predictive variance, enabling analytic propagation of predictive variance through the linear decoder without Monte Carlo estimation in pixel space. Across sea surface temperature, medical imaging, and natural images, the method achieves strong reconstruction with fewer parameters and lower memory, while producing well-calibrated spatial uncertainty that closely matches empirical ensembles.

超分辨率扩散模型不确定性量化POD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。