不同估算方法让JS散度值不可比,论文提出校正方案与工具。
Not all Jensen-Shannon Divergence Estimators are Equal

- 用分类器法估算联合分布结构,避免边缘法忽略依赖
- 在高维或类别不平衡时,估计值偏差显著,最大达30%以上
- 适合数据合成评估者、模型对比研究者使用
JS散度被广泛用于衡量合成表格数据的质量,但实际中常基于有限样本估算,而估算协议常不明确,导致测量结果不可靠。尽管总体散度有明确定义,但经验值受估算方法、采样方式、校准、维度和类别平衡影响。我们发现:基于边缘的估算方法会忽略联合分布依赖,严重低估散度;基于分类器的方法虽能捕捉联合结构,却对估算器高度敏感。通过控制实验与真实世界合成表基准测试,我们揭示了边缘估算的依赖盲区、类别不平衡下的先验偏移问题及高维下的估算敏感性。为解决先验偏移,我们推导出分类器法的闭式后验校正公式。结果表明,经验JS散度值本质上依赖于估算协议,必须明确定义才能进行有效比较。本文提供实用指南与开源工具,支持估算感知的JS散度评估。
原文摘要 · Abstract (English)
The Jensen-Shannon divergence is widely reported as a scalar measure of fidelity for synthetic tabular data. Yet, in practice, it is estimated from finite samples using protocols that are often underspecified. This creates a measurement problem. Although the population divergence is well defined, the empirical value depends on the estimator family, sampling protocol, calibration, dimensionality, and class balance. We show that different protocols can yield non-comparable values: marginal-based estimators ignore dependencies in the joint distribution and can severely underestimate divergence, while classifier-based estimators capture joint structure but exhibit strong estimator dependence. We systematically study this behavior across controlled settings with reference divergences and real-world synthetic tabular benchmarks. Our analysis reveals dependence blindness in marginal estimators, prior-shift bias under class imbalance, and estimator sensitivity in high dimensions. To address prior shift, we derive a closed-form posterior correction for classifier-based Jensen-Shannon estimation. Our results show that empirical Jensen-Shannon divergence values are inherently protocol-dependent, making explicit specification of the estimation procedure necessary for meaningful comparison. We provide practical guidelines and an open-source tool for estimator-aware Jensen-Shannon evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。