arXiv:2607.04977cs.LGstat.ML2026-07

基于几何感知的贝叶斯量化方法,提升标签分布估计精度

Geometry-Aware Bayesian Quantification via Compositional Data Analysis

论文配图:Geometry-Aware Bayesian Quantification via Compositional Data Analysis
图 1 · 摘自论文原文
  • 用对数比表示和艾奇森几何建模后验向量的单纯形结构
  • 在42个数据集上优于标准KDE基线,贝叶斯版本表现更强
  • 适合需要可靠置信度的标签偏移场景,如医疗、金融分类

准确估计未知目标标签分布是应对标签偏移的关键第一步。该任务被称为量化或类别先验估计,近年通过基于连续核密度估计(KDE)的方法取得显著进展,这些方法建模多类分类器后验概率的密度。由于后验向量位于概率单纯形上,可视为组合数据。然而,现有KDE量化方法通常使用欧氏高斯核,忽略单纯形几何结构,错误地将概率质量分配到边界外。本文提出一种基于对数比表示与艾奇森几何的几何感知KDE模型,并引入收缩正则化以增强边界附近的鲁棒性。结合KDE量化最大似然解释,推导出类别先验的点估计与贝叶斯推断方法。在涵盖表格、文本和图像领域的42个数据集上的实验表明,该方法在性能上媲美最先进量化器,常优于标准KDE基线,且在贝叶斯量化方法中表现优异。

原文摘要 · Abstract (English)

Accurately estimating the unknown target label distribution is the critical first step for adapting to label shift. This task, widely known as quantification or class prevalence estimation, has recently seen significant advances through continuous KDE-based methods which model the density of multiclass classifier posteriors. Posterior vectors might be regarded as compositional data, since they lie on the probability simplex. However, existing KDE-based quantifiers typically rely on Euclidean Gaussian kernels, which ignore simplex geometry and incorrectly assign probability mass outside its boundaries. We introduce a geometry-aware KDE model for multiclass quantification based on log-ratio representations and Aitchison geometry, together with a shrinkage regularization that improves robustness near the simplex boundary. Combined with a maximum-likelihood interpretation of KDE-based quantification, we derive both point-estimation and Bayesian inference procedures for class prevalences. Experiments on 42 datasets across tabular, text, and image domains show that the proposed method is competitive with state-of-the-art quantifiers, often improving over standard KDE-based baselines, while also yielding strong results among Bayesian quantification methods.

量化贝叶斯几何建模标签偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。