arXiv:2605.06413stat.MLcs.LG2026-05

分离不确定性类型,提升贝叶斯优化在噪声环境下的决策能力

Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors

论文配图:Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors
图 1 · 摘自论文原文
  • 通过可控合成数据构建信号与噪声解耦的双头模型
  • 在异方差噪声下,仅基于认知不确定性采样可避免盲目探索
  • 适合需要精准不确定性建模的超参优化与主动学习场景

先验拟合网络(PFNs)通过元学习合成任务先验来近似贝叶斯预测,但其标准输出为观测噪声下的后验预测分布。在序列决策任务中,如主动学习与贝叶斯优化,应优先关注对潜在信号的认知不确定性,而非不可约的随机噪声。我们证明,仅从后验预测分布无法识别这种不确定性分解,即使分布已知。利用PFNs的优势——可控制合成数据生成过程,每个任务显式包含潜在信号和噪声函数,并提供查询级别的无噪目标与观测噪声方差标签。我们训练一个解耦的PFN,具有独立的信号与异方差头。观测级预测由潜信号分布与学习到的噪声模型卷积得到。实验表明,仅使用认知不确定性进行采样可缓解总方差探索的失败模式。在匹配对比中,解耦模型通常优于调优后的观测级基线,尤其在超参数优化(HPO)中提升显著;在更广泛的测试中,该模型在HPO与合成贝叶斯优化中均取得最佳平均排名。

原文摘要 · Abstract (English)

Prior-Fitted Networks (PFNs) amortize Bayesian prediction by meta-learning over a synthetic task prior, but their standard output is a posterior predictive distribution over noisy observations. For sequential decision-making, such as active learning and Bayesian optimization, acquisition should prioritize epistemic uncertainty about the latent signal rather than irreducible aleatoric observation noise. We show that this epistemic--aleatoric split is not identifiable in general from the posterior predictive distribution alone, even when that distribution is known exactly. We then exploit a distinctive advantage of PFNs: because the synthetic data-generating process is under our control, each task can contain an explicit latent signal and noise function, and the generator can provide query-level labels for both the noiseless target and the observation-noise variance. We use these labels to train a decoupled PFN with separate latent-signal and aleatoric heads. The observation-level predictive is induced by convolving the latent signal distribution with the learned noise model. Empirically, epistemic-only acquisition mitigates the failure mode of total-variance exploration in noisy and heteroscedastic settings. In matched comparisons, decoupled models usually improve over tuned observation-level baselines, with the clearest gains in HPO; in broader sweeps, a decoupled model obtains the best average rank in both HPO and synthetic BO.

不确定性量化贝叶斯优化元学习解耦表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。