构建复杂数据下的互信息评估新基准,揭示无绝对最优估计器。
Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation
- 从耦合理论出发,设计两类测试:先生成耦合关系、再控制边缘分布。
- 在真实图像与合成数据上验证,三类估计算法各有优势。
- 发现不同场景下性能差异大,适合研究互信息估计的学者参考。
互信息(MI)估计是机器学习与统计学的核心问题,但现有基准多基于简化、低维分布,难以反映算法在复杂真实数据上的表现。本文提出一个统一的耦合理论框架,涵盖已有基准作为特例,并设计两类互补测试:一类以耦合为先,通过合成与流模型变换系统调节真实MI值、维度与边缘复杂度;另一类以边缘为先,将真实图像数据与可控依赖结构配对,扩展经典同类别配对范式。在此套件中,我们全面评估了三类估计算法:非参数型、判别型与生成型。结果表明,不存在普遍最优的估计器:各类方法在特定条件下可显著超越其他类型。通过对这些情形的分析,我们识别出关键估计障碍,并提出能更有效暴露其局限性的新测试。代码已开源于 https://github.com/VanessB/mutinfo。
原文摘要 · Abstract (English)
Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate estimators on simplified, low-dimensional distributions, leaving their performance on complex, realistic data largely unexplored. We address this gap with a comprehensive benchmarking framework grounded in a unified copula-theoretic perspective that subsumes existing benchmarks as special cases. Within this framework, we propose two complementary families of tests: a copula-first family that systematically varies ground-truth MI, dimensionality, and marginal complexity using synthetic and flow-based transformations; and a marginals-first family that couples real-world image data with controlled dependency structures, extending the classic same-class-pairing paradigm. We use this suite to extensively evaluate three classes of estimators: non-parametric, discriminative, and generative. Contrary to prevailing assumptions, our results indicate that there is no universal winner: each category can systematically outperform all other estimators under specific setups. By analyzing these cases, we identify fundamental estimation barriers and propose new tests that more effectively stress these specific limitations. We share the open source code at https://github.com/VanessB/mutinfo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。