arXiv:2601.10641stat.MEcs.LG2026-01

提出新方法解决相似性度量在特定场景下失效问题

Adjusted Similarity Measures and a Violation of Expectations

  • 推广调整算子至多种零假设模型,支持更灵活的评估框架
  • 发现传统调整可能导致度量为负或恒为0,严重偏离预期均值0
  • 适用于聚类评估、一致性检验等需可靠相似性度量的场景

调整后的相似性度量(如Cohen's kappa和调整兰德指数)是评估离散标签一致性的关键工具,其理想性质是在零假设下期望值为0,最大相似时取值为1。当前常基于置换分布进行调整,但近年对更符合实际情境的零模型(如允许随机簇数的聚类集成)的关注日益增加。本文旨在:(1) 将调整算子推广至一般零模型,并包含统计标准化作为特例;(2) 识别保证理想性质的充分条件,这些条件与观测数据是否纳入零分布相关。我们证明,若不满足这些条件,可能引发严重问题:传统调整导致度量非正而非均值0,或统计标准化下度量恒为0。

原文摘要 · Abstract (English)

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the property of 0 expectation under a null distribution and maximum value 1 under maximal similarity to aid in interpretation. Measures are frequently adjusted with respect to the permutation distribution for historic and analytic reasons. There is currently renewed interest in considering other null models more appropriate for context, such as clustering ensembles permitting a random number of identified clusters. The purpose of this work is two -- fold: (1) to generalize the study of the adjustment operator to general null models and to a more general procedure which includes statistical standardization as a special case and (2) to identify sufficient conditions for the adjustment operator to produce the intended properties, where sufficient conditions are related to whether and how observed data are incorporated into null distributions. We demonstrate how violations of the sufficient conditions may lead to substantial breakdown, such as by producing a non-positive measure under traditional adjustment rather than one with mean 0, or by producing a measure which is deterministically 0 under statistical standardization.

相似性度量聚类评估统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。