厘清记忆化在可信机器学习中的三重粒度,揭示其利弊权衡。
Trustworthy Machine Learning via Memorization and the Granular Long-Tail: A Survey on Interactions, Tradeoffs, and Beyond
- 提出类不平衡、异常性与噪声的三重粒度框架
- 发现现有方法混淆异常样本与噪声导致误判
- 为公平性、鲁棒性与隐私保护提供新研究路径
机器学习中的记忆化现象日益受到关注,现代模型被实证发现会记忆训练数据片段。以往理论分析(如Feldman的工作)指出,长尾分布是记忆化的根源,且对尾部样本不可避免。然而,记忆化与可信机器学习的交叉研究暴露出关键空白:现有工作仅关注类别不平衡,而近期研究开始区分类别罕见性与同类中稀有但有效的异常样本。当前框架却将异常样本与噪声、错误数据混为一谈,忽略了它们在公平性、鲁棒性与隐私上的不同影响。本文系统综述了可信机器学习与记忆化相关的研究成果,揭示未被探索的缺口,并提出新的研究方向。由于现有理论与实证分析缺乏区分记忆化双重性的精细度,我们形式化提出三重长尾粒度——类别不平衡、异常性、噪声,以揭示现有框架如何误用这些层次,从而持续产生缺陷解决方案。通过系统化该粒度,我们绘制出未来研究路线图。可信机器学习必须调和记忆异常性以保障公平性,与抑制噪声以确保鲁棒性和隐私之间的微妙权衡。通过此粒度重新定义记忆化,重塑可信机器学习的理论基础,并为实现性能与社会信任对齐的模型提供实证前提。
原文摘要 · Abstract (English)
The role of memorization in machine learning (ML) has garnered significant attention, particularly as modern models are empirically observed to memorize fragments of training data. Previous theoretical analyses, such as Feldman's seminal work, attribute memorization to the prevalence of long-tail distributions in training data, proving it unavoidable for samples that lie in the tail of the distribution. However, the intersection of memorization and trustworthy ML research reveals critical gaps. While prior research in memorization in trustworthy ML has solely focused on class imbalance, recent work starts to differentiate class-level rarity from atypical samples, which are valid and rare intra-class instances. However, a critical research gap remains: current frameworks conflate atypical samples with noisy and erroneous data, neglecting their divergent impacts on fairness, robustness, and privacy. In this work, we conduct a thorough survey of existing research and their findings on trustworthy ML and the role of memorization. More and beyond, we identify and highlight uncharted gaps and propose new revenues in this research direction. Since existing theoretical and empirical analyses lack the nuances to disentangle memorization's duality as both a necessity and a liability, we formalize three-level long-tail granularity - class imbalance, atypicality, and noise - to reveal how current frameworks misapply these levels, perpetuating flawed solutions. By systematizing this granularity, we draw a roadmap for future research. Trustworthy ML must reconcile the nuanced trade-offs between memorizing atypicality for fairness assurance and suppressing noise for robustness and privacy guarantee. Redefining memorization via this granularity reshapes the theoretical foundation for trustworthy ML, and further affords an empirical prerequisite for models that align performance with societal trust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。