arXiv:2511.03849cs.ITcs.LG2025-11

比较两种相似性敏感熵的优劣,发现多数情况下应选LCR。

Which Similarity-Sensitive Entropy (Sentropy)?

  • 用半距离参数化相似度缩放,分析LCR与VS差异
  • 53个数据集实验显示两者数值可差几个数量级
  • 除特殊情形外,一般应优先使用LCR方法

香农熵并非机器学习数据集中唯一相关或最重要的熵,传统熵仅捕捉元素频率信息,而忽略其相似性与差异性所编码的丰富信息。这需要引入相似性敏感熵(称作“sentropy”)。Sentropy可通过最近的Leinster-Cobbold-Reeve框架(LCR)或更新的Vendi分数(VS)测量。本文通过53个大型图像与表格数据集,从理论和数值两方面探讨该选择问题。结果表明,LCR与VS值可相差数个数量级,二者互补,但非普遍适用。我们证明了对于所有非负瑞尼-希尔阶参数,以及在相似度矩阵满秩时负值情况,VS始终为LCR的上界。结论是:仅当数据元素可解释为基本‘原元素’的线性组合,或系统具有量子力学特征时,才推荐使用VS;否则,在一般情况下,建议采用LCR以更全面捕捉相似性、差异性和频率信息,特定半距离下二者可互补。

原文摘要 · Abstract (English)

Shannon entropy is not the only entropy that is relevant to machine-learning datasets, nor possibly even the most important one. Traditional entropies such as Shannon entropy capture information represented by elements' frequencies but not the richer information encoded by their similarities and differences. Capturing the latter requires similarity-sensitive entropy (``sentropy''). Sentropy can be measured using either the recently developed Leinster-Cobbold-Reeve framework (LCR) or the newer Vendi score (VS). This raises the practical question of which one to use: LCR or VS. Here we address this question theoretically and numerically, using 53 large and well-known imaging and tabular datasets. We find that LCR and VS values can differ by orders of magnitude and are complementary, except in limiting cases. We show that both LCR and VS results depend on how similarities are scaled, and introduce the notion of ``half-distance'' to parameterize this dependence. We prove the VS provides an upper bound on LCR for all non-negative values of the Rényi-Hill order parameter, as well as for negative values in the special case that the similarity matrix is full rank. We conclude that VS is preferable only when a dataset's elements can be usefully interpreted as linear combinations of a more fundamental set of ``ur-elements'' or when the system that the dataset describes has a quantum-mechanical character. In the broader case where one simply wishes to capture the rich information encoded by elements' similarities and differences as well as their frequencies, we propose that LCR should be favored; nevertheless, for certain half-distances the two methods can complement each other.

相似性数据分析机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。