模型影响数据分布时,正则化可提前应对这种动态变化。
Optimal Regularization for Performative Learning
- 用高维岭回归分析正则化如何缓解模型引发的数据变化
- 过参数化时,性能扰动反而能降低测试风险
- 最优正则化强度与扰动整体强度成正比,适合动态系统
在绩效学习中,部署的模型会改变数据分布——例如策略性用户调整自身特征以操纵模型结果——这使得其动态比经典监督学习更复杂。因此,不仅需针对当前数据优化模型,还需考虑模型可能引导数据分布转向新方向,但又无法预知具体变化。本文研究正则化在高维岭回归中对绩效效应的应对作用。结果显示,尽管绩效效应在总体设置下会增加测试风险,但在特征数超过样本数的过参数化情形中,它反而可能带来益处。我们证明最优正则化强度与整体绩效效应强度成正比,从而可在部署前预设正则化参数以应对潜在扰动。通过合成与真实数据集的实证评估,验证了该发现的有效性。
原文摘要 · Abstract (English)
In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it - which creates a more complex dynamic than in classical supervised learning. One should thus not only optimize the model for the current data but also take into account that the model might steer the distribution in a new direction, without knowing the exact nature of the potential shift. We explore how regularization can help cope with performative effects by studying its impact in high-dimensional ridge regression. We show that, while performative effects worsen the test risk in the population setting, they can be beneficial in the over-parameterized regime where the number of features exceeds the number of samples. We show that the optimal regularization scales with the overall strength of the performative effect, making it possible to set the regularization in anticipation of this effect. We illustrate this finding through empirical evaluations of the optimal regularization parameter on both synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。