在算法影响数据分布时,如何设计更鲁棒的分类模型?
Optimal Classification under Performative Distribution Shift
- 用映射变换建模算法决策带来的数据分布变化
- 证明线性参数分类下风险函数为凸,可高效优化
- 适用于真实数据集,对强干扰方向也有效
当算法决策公开部署后可能引起数据分布变化时,性能化学习(performative learning)提供了应对策略。本文提出将这种影响建模为前推测度(push-forward measures),构建通用框架,支持新型梯度估计方法,提升学习效率与可扩展性。与以往需完全知晓分布不同,本方法仅需掌握描述变化的转移算子。该框架可集成至基于变量变换的模型,如变分自编码器(VAEs)或归一化流(normalizing flows)。聚焦于线性参数的性能化效应分类问题,我们证明在新假设下性能化风险具有凸性。关键在于不约束性能效应强度,而限定其方向:部署更精准模型会导致分类难度上升。此时,性能化风险最小化可重述为一个极小极大变分问题,与对抗鲁棒分类建立联系。实验在合成与真实数据集上验证了方法有效性。
原文摘要 · Abstract (English)
Performative learning addresses the increasingly pervasive situations in which algorithmic decisions may induce changes in the data distribution as a consequence of their public deployment. We propose a novel view in which these performative effects are modelled as push-forward measures. This general framework encompasses existing models and enables novel performative gradient estimation methods, leading to more efficient and scalable learning strategies. For distribution shifts, unlike previous models which require full specification of the data distribution, we only assume knowledge of the shift operator that represents the performative changes. This approach can also be integrated into various change-of-variablebased models, such as VAEs or normalizing flows. Focusing on classification with a linear-in-parameters performative effect, we prove the convexity of the performative risk under a new set of assumptions. Notably, we do not limit the strength of performative effects but rather their direction, requiring only that classification becomes harder when deploying more accurate models. In this case, we also establish a connection with adversarially robust classification by reformulating the minimization of the performative risk as a min-max variational problem. Finally, we illustrate our approach on synthetic and real datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。