arXiv:2510.08335stat.MLcs.LG2025-10

解决机器学习模型部署后数据分布变化导致性能下降的问题

PAC Learnability in the Presence of Performativity

  • 构建仅依赖原始数据的可偏倚风险估计函数
  • 证明标准可学习假设空间在扰动场景下仍可学习
  • 适用于真实与合成数据,适合部署后性能优化场景

随着机器学习模型在现实应用中的广泛采用,模型引发的数据分布变化(即绩效性)日益普遍。由于模型通常仅基于原始(未扰动)分布训练,这种绩效性变化可能导致测试时性能下降。本文从经典的PAC(可能近似正确)学习框架出发,研究绩效性二分类问题是否可学习。我们提出了多种绩效性场景,包括标签分布的线性偏移及特征与标签的更一般变化。构造了一种仅依赖原始分布数据和绩效效应类型的绩效性经验风险函数,该函数对扰动分布上的真实风险是无偏估计。最小化此风险可证明:在标准二分类中可学习的假设空间,在所考虑的绩效性场景下仍保持可学习性。我们还对绩效性风险最小化方法进行了广泛的实验评估,并在合成与真实数据上展示了其优势。

原文摘要 · Abstract (English)

Following the wide-spread adoption of machine learning models in real-world applications, the phenomenon of performativity, i.e. model-dependent shifts in the test distribution, becomes increasingly prevalent. Unfortunately, since models are usually trained solely based on samples from the original (unshifted) distribution, this performative shift may lead to decreased test-time performance. In this paper, we study the question of whether and when performative binary classification problems are learnable, via the lens of the classic PAC (Probably Approximately Correct) learning framework. We motivate several performative scenarios, accounting in particular for linear shifts in the label distribution, as well as for more general changes in both the labels and the features. We construct a performative empirical risk function, which depends only on data from the original distribution and on the type performative effect, and is yet an unbiased estimate of the true risk of a classifier on the shifted distribution. Minimizing this notion of performative risk allows us to show that any PAC-learnable hypothesis space in the standard binary classification setting remains PAC-learnable for the considered performative scenarios. We also conduct an extensive experimental evaluation of our performative risk minimization method and showcase benefits on synthetic and real data.

PAC学习绩效性分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。