arXiv:2501.10815cs.LGmath.ST2025-01中稿 · 2025 SIAM Internat…

新度量方法可解释地量化连续变量间的预测依赖关系。

An Interpretable Measure for Quantifying Predictive Dependence between Continuous Random Variables -- Extended Version

  • 基于预测损失的非参数度量,捕捉多种复杂关系。
  • 在9万+数据集上优于现有方法,能发现传统方法忽略的依赖。
  • 结果可解释性强,适合需要理解变量关联的研究者。

统计学习中的基础任务是量化两个连续随机变量之间的联合依赖或关联。本文提出一种全新的全非参数度量,用于评估连续变量 $X$ 与 $Y$ 之间的关联程度,能够捕捉包括非函数关系在内的广泛关系类型。该度量的核心优势在于可解释性:它量化了在预测 $Y$ 时忽略 $X$ 分布所导致的预测准确率的期望相对损失。该度量取值范围为 [0,1],且仅当 $X$ 与 $Y$ 独立时等于零。我们在超过90,000个真实和合成数据集上评估了该度量性能,并与主流方法进行对比。结果表明,该度量能提供有价值的潜在关系洞察,尤其在现有方法无法捕捉重要依赖的情况下表现更优。

原文摘要 · Abstract (English)

A fundamental task in statistical learning is quantifying the joint dependence or association between two continuous random variables. We introduce a novel, fully non-parametric measure that assesses the degree of association between continuous variables $X$ and $Y$, capable of capturing a wide range of relationships, including non-functional ones. A key advantage of this measure is its interpretability: it quantifies the expected relative loss in predictive accuracy when the distribution of $X$ is ignored in predicting $Y$. This measure is bounded within the interval [0,1] and is equal to zero if and only if $X$ and $Y$ are independent. We evaluate the performance of our measure on over 90,000 real and synthetic datasets, benchmarking it against leading alternatives. Our results demonstrate that the proposed measure provides valuable insights into underlying relationships, particularly in cases where existing methods fail to capture important dependencies.

统计学习依赖度量可解释性非参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。