arXiv:2604.06378cs.GTcs.LG2026-04

当人们会因算法分类调整行为时,公平性难题可被破解,但需差异化对待相同结果。

Revisiting Fairness Impossibility with Endogenous Behavior

  • 让算法先统一各群体统计表现,再调节奖惩力度以引导相似行为。
  • 在策略性行为下,错误率平衡与预测准确性不再冲突,但可能产生新不公。
  • 适合关注算法公平性与人类行为互动的政策设计者阅读。

在许多现实场景中,机构可调整算法分类决策的后果,如罚款金额、判刑长度或福利水平,这些后果称为分类的赌注。赌注可能引发人们对分类的策略性反应,即人们会根据预期分类结果调整自身行为。现有算法公平性研究多假设行为固定,将群体间差异视为环境外生特征,忽略赌注对结果的影响。本文重新审视此类设定下的经典不可能性结果,发现当个体可策略性响应分类时,错误率平衡与预测准确性的不兼容性消失,但仅通过在相同分类结果下对不同群体施加不同后果来实现。我们提出两阶段设计:先确保各群体统计性能一致,再调节赌注以诱导相似行为模式。这要求对不同群体区别对待相同决策的后果。结果表明,战略环境下公平性不能仅由算法映射关系决定,必须将人类后果作为核心设计变量,引入规范准则,并揭示其与统计公平标准的新型权衡。本文旨在明确并显式化这些权衡。

原文摘要 · Abstract (English)

In many real-world settings, institutions can and do adjust the consequences attached to algorithmic classification decisions, such as the size of fines, sentence lengths, or benefit levels. We refer to these consequences as the stakes associated with classification. These stakes can give rise to behavioral responses to classification, as people adjust their actions in anticipation of how they will be classified. Much of the algorithmic fairness literature evaluates classification outcomes while holding behavior fixed, treating behavioral differences across groups as exogenous features of the environment. Under this assumption, the stakes of classification play no role in shaping outcomes. We revisit classic impossibility results in algorithmic fairness in a setting where people respond strategically to classification. We show that, in this environment, the well-known incompatibility between error-rate balance and predictive parity disappears, but only by potentially introducing a qualitatively different form of unequal treatment. Concretely, we construct a two-stage design in which a classifier first standardizes its statistical performance across groups, and then adjusts stakes so as to induce comparable patterns of behavior. This requires treating groups differently in the consequences attached to identical classification decisions. Our results demonstrate that fairness in strategic settings cannot be assessed solely by how algorithms map data into decisions. Rather, our analysis treats the human consequences of classification as primary design variables, introduces normative criteria governing their use, and shows that their interaction with statistical fairness criteria generates qualitatively new tradeoffs. Our aim is to make these tradeoffs precise and explicit.

算法公平策略行为公平性权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。