用因果特征提升分类鲁棒性,让机构用户长期目标一致
The Role of Causal Features in Strategic Classification for Robustness and Alignment

- 基于因果模型设计分类器,应对用户主动改变特征带来的分布漂移
- 在噪声有界条件下,因果分类可实现最优误差,且能分解出偏差来源
- 实验证明理论预测有效,适合关注公平与长期激励对齐的研究者
在策略分类中,机构(如银行)需预判用户为提高分类结果(如贷款通过率)而主动调整自身特征的行为。由于用户行为导致的分布漂移是核心挑战,本文引入因果模型,因其可界定最坏情况下的分布外(OOD)风险。首先,当噪声满足特定有界条件时,因果分类可实现任意充分大适应后的最优分类误差。其次,若假设不成立,最优分类器的OOD交叉熵风险可分解为一个OOD偏差项和一个因未使用全部可观测特征引起的项,从而揭示因果分类的优势场景。最后,我们证明使用因果特征能实现机构与用户长期激励的一致,区别于以往认为此类方法带来社会成本的观点。我们在合成数据上验证了理论,结果显示理论预测与实际行为高度一致。
原文摘要 · Abstract (English)
In strategic classification, an institution (e.g., a bank) anticipates adaptation from users who change their features to increase utility in a classification task (e.g., loan repayment). Since a key challenge is the distribution shift induced by users, we turn to causal models, which have been shown to bound the worst-case out-of-distribution (OOD) risk, and establish several new results that link causality and strategic classification. First, we show that causal classification leads to optimal classification error after any sufficiently large adaptation, when the noise is bounded in a certain way. Second, when these assumptions do not hold, we show OOD cross-entropy risk of optimal classifiers decomposes into an OOD bias term and a term arising from not using all observable features, allowing us to understand when causal classifiers have an advantage. Finally, we show that the use of causal features can allow alignment of long-term incentives between institutions and users, contrasting with previous work that highlights social costs of such approaches. We validate our theory empirically on synthetic data, finding that our results predict behavior in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。