arXiv:2508.02337cs.LGstat.CO2025-08

提出可扩展的后验采样方法,更好估计词向量不确定性。

Posterior Sampling of Probabilistic Word Embeddings

  • 用多项式-伽马增强和拉普拉斯近似构建可扩展吉布斯采样器
  • 在真实数据上验证方法可行,小样本下后验均值优于最大后验估计
  • 解决词向量非唯一性问题,适合需要可靠不确定性的场景

量化词向量中的不确定性对文本数据的可靠推断至关重要。现有贝叶斯方法如哈密顿蒙特卡洛(HMC)和均值场变分推断(MFVI)或计算成本过高,或依赖严格假设。本文提出基于多项式-伽马增强的可扩展吉布斯采样器,结合拉普拉斯近似,并与MFVI和HMC对比词向量的不确定性估计。我们解决了词向量的非唯一性问题。实验表明,吉布斯采样器和HMC能正确估计不确定性,而MFVI不能,拉普拉斯近似仅在大样本时有效。将吉布斯采样器应用于美国国会和Movielens数据集,验证其在大规模真实数据上的可行性。由于能获取完整后验样本,后验均值在保留样本似然上优于最大后验(MAP)估计,尤其在小样本时更显著,进一步证明词向量后验采样的必要性。

原文摘要 · Abstract (English)

Quantifying uncertainty in word embeddings is crucial for reliable inference from textual data. However, existing Bayesian methods such as Hamiltonian Monte Carlo (HMC) and mean-field variational inference (MFVI) are either computationally infeasible for large data or rely on restrictive assumptions. We propose a scalable Gibbs sampler using Polya-Gamma augmentation as well as Laplace approximation and compare them with MFVI and HMC for word embeddings. In addition, we address non-identifiability in word embeddings. Our Gibbs sampler and HMC correctly estimate uncertainties, while MFVI does not, and Laplace approximation only does so on large sample sizes, as expected. Applying the Gibbs sampler to the US Congress and the Movielens datasets, we demonstrate the feasibility on larger real data. Finally, as a result of having draws from the full posterior, we show that the posterior mean of word embeddings improves over maximum a posteriori (MAP) estimates in terms of hold-out likelihood, especially for smaller sampling sizes, further strengthening the need for posterior sampling of word embeddings.

词向量不确定性后验采样贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。