arXiv:2608.10398cs.LGcs.AI2026-08

ELVAE通过证据学习实现生成不确定性感知,分离出可调控的敏感性机制。

ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation

论文配图:ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation
图 1 · 摘自论文原文
  • 在变分自编码器潜空间引入输入依赖的正逆伽马先验,解耦位置不确定性和条件变异性。
  • 训练后逆证据1/ν为最强敏感性指标,在MNIST上高/低不确定性语义转换率可达1.98。
  • 模型暴露可控敏感性机制,适合需要生成可信度评估的AI应用,如医疗影像生成。

ELVAE在每个VAE潜变量坐标上设置输入相关的正逆伽马(NIG)层级,将位置不确定性 $u_{\mathrm{epi}}=β/[ν(α-1)]$ 与条件变异性 $u_{\mathrm{var}}=β/(α-1)$ 分离。然而,边缘化潜变量分布仅识别三个商坐标 $(γ,α,c)$,其中 $c=β(1+1/ν)$;重建过程对一个 $(ν,β)$ 纤维方向完全不可见。理论分析表明,完整的NIG先验与前向KL选择每根纤维上的唯一相对规范代表,因此规范逆分配并非第四个独立信息通道。实证中,训练后的逆证据 $1/ν$ 始终是最重要的敏感性排序得分。当 $τ_{\mathrm{epi}}=1$ 时,三种子均值的高/低 $u_{\mathrm{epi}}$ 语义转换比率在MNIST上为1.98,在Fashion-MNIST上为1.66,经尺度匹配控制后降至1.33和1.16。在20次抽样的MNIST组件研究中,等锚点扰动能量下,$1/ν$ 在 $u_{\mathrm{epi}}$ 和 $u_{\mathrm{var}}$ 场中分别给出1.69和1.65的高/低比率,在无几何约束的各向同性场中为1.69(95%区间1.45–1.97),而 $u_{\mathrm{var}}$ 将各向同性顺序反转至0.76。在实验先验下,规范逆证据 $1/ν_{\mathrm{can}}$ 是 $T=c/[α(γ^2+2)]$ 的严格单调变换,且上界为 $3+\sqrt{10}$。因此,ELVAE揭示了一个可控的敏感性机制,其训练后的四输出实现具有操作意义,而可见重建信息仍保持三维,基础图像质量则为独立问题。

原文摘要 · Abstract (English)

ELVAE places an input-dependent normal--inverse-gamma (NIG) hierarchy at each VAE latent coordinate, separating location uncertainty $u_{\mathrm{epi}}=β/[ν(α-1)]$ from conditional variability $u_{\mathrm{var}}=β/(α-1)$. The marginalized latent law, however, identifies only the three quotient coordinates $(γ,α,c)$ with $c=β(1+1/ν)$; reconstruction is blind to one $(ν,β)$ fiber direction. A companion theoretical analysis shows that the complete NIG prior and forward KL select a unique prior-relative canonical representative on each fiber, so canonical inverse allocation is not a fourth independent information channel. Empirically, trained inverse evidence $1/ν$ remains the strongest sensitivity-ranking score. At $τ_{\mathrm{epi}}=1$, the three-seed mean high/low-$u_{\mathrm{epi}}$ semantic-transition ratios are 1.98 on MNIST and 1.66 on Fashion-MNIST, falling to 1.33 and 1.16 under scale-matched controls. In the 20-draw MNIST component study with equal per-anchor perturbation energy, $1/ν$ gives high/low ratios 1.69 and 1.65 under the $u_{\mathrm{epi}}$ and $u_{\mathrm{var}}$ fields and 1.69 (95\% interval 1.45--1.97) under a geometry-free isotropic field, whereas $u_{\mathrm{var}}$ reverses the isotropic ordering to 0.76. Under the experimental prior, canonical $1/ν_{\mathrm{can}}$ is a strictly increasing transform of $T=c/[α(γ^2+2)]$ and is bounded above by $3+\sqrt{10}$. Thus ELVAE exposes a controllable sensitivity mechanism whose trained four-output realization is operationally informative, while the exact reconstruction-visible information remains three-dimensional and baseline image quality is a separate question.

生成模型不确定性变分自编码器证据学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。