arXiv:2411.04696cs.LGcs.AI2024-11

机器学习中的虚假相关性不是单纯统计问题,而是由实际影响决定的判断。

The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter

  • 用四个实用框架评估相关性是否该用:任务相关、泛化能力、类人逻辑、无害性
  • 虚假相关性的判定依赖模型表现和伦理后果,而非固定统计标准
  • 适合关注公平性、鲁棒性和可解释性的研究者阅读

机器学习依赖数据中发现的相关性,但模型常捕获无意间形成的虚假相关,导致性能下降、偏见加剧。本文追溯虚假相关性在统计学中的传统定义(非因果关联),发现当前机器学习研究中其含义已被重新诠释。研究者不依赖形式定义,而是通过‘实用框架’判断相关性是否合理:任务相关性(相关性应服务于任务)、泛化性(能在未见数据上成立)、类人性(符合人类认知方式)、无害性(不引发社会或伦理问题)。这些框架表明,相关性的优劣并非固定属性,而是由技术、认知与伦理因素共同决定的实践判断。通过对这一核心难题的文献分析,本文揭示了技术概念如何在实践中被动态建构,推动对虚假相关性等关键概念的深层理解。

原文摘要 · Abstract (English)

Learning correlations from data forms the foundation of today's machine learning (ML) and artificial intelligence research. While contemporary methods enable the automatic discovery of complex patterns, they are prone to failure when unintended correlations are captured. This vulnerability has spurred a growing interest in interrogating spuriousness, which is often seen as a threat to model performance, fairness, and robustness. In this article, we trace departures from the conventional statistical definition of spuriousness-which denotes a non-causal relationship arising from coincidence or confounding-to examine how its meaning is negotiated in ML research. Rather than relying solely on formal definitions, researchers assess spuriousness through what we call pragmatic frames: Judgments based on what a correlation does in practice-how it affects model behavior, supports or impedes task performance, or aligns with broader normative goals. Drawing on a broad survey of ML literature, we identify four such frames: Relevance (Models should use correlations that are relevant to the task), generalizability (Models should use correlations that generalize to unseen data), human-likeness (Models should use correlations that a human would use to perform the same task), and harmfulness (Models should use correlations that are not socially or ethically harmful). These representations reveal that correlation desirability is not a fixed statistical property but a situated judgment informed by technical, epistemic, and ethical considerations. By examining how a foundational ML conundrum is problematized in research literature, we contribute to broader conversations on the contingent practices through which technical concepts like spuriousness are defined and operationalized.

虚假相关可解释性伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。