通过多阶段选择性分类器,激励用户长期努力而非作弊。
Sequential Strategic Classification with Multi-Stage Selective Classifiers

- 设计多阶段分类系统,允许分类器在不确定时放弃判断。
- 证明长期努力策略优于短期作弊,在多轮中收益更高。
- 适合关注公平决策与行为激励的系统设计者。
传统战略分类研究关注个体为获取有利结果而操纵特征,通常导致欺骗行为。现有工作多聚焦单次或重复使用同一分类器的场景,但现实决策常为多阶段、逐步升级的过程。本文提出一种序列化、随机的多阶段战略分类模型,模拟个体在不同阶段通过提升可观测特征与真实属性(改进)或仅提升可观测特征(作弊)来适应难度递增的分类任务。每阶段采用可拒绝预测的选择性分类器:成功则晋级,失败则降级,拒绝则停留原级。我们完整刻画了在选择性分类器下个体的最优即时行为,并比较了长期遵循无改进(从不改进)或无作弊(从不作弊)的短视策略的效用。进一步分析分类器序列的设计原则,发现特定设计能使长期努力策略更具优势,从而有效激励真实努力。
原文摘要 · Abstract (English)
Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes made by classifiers, typically turning to dishonest actions when they are less costly than genuine efforts. Prior works have demonstrated a fundamental inability to get out of this conundrum by only focusing on the design of a classifier. We note that prior work also heavily focuses on either one-shot settings or repeated interaction with the same classifier. Real-world decision making is often multi-stage, involving a sequence of potentially different classifiers as an agent progresses. This paper introduces a sequential, stochastic, multi-stage model of strategic classification, by capturing how agents adapt their behavior, through improvement actions (enhancing both observable features and true attributes) and gaming actions (enhancing only observable features), over multiple levels of classification with increasing difficulty as well as reward. For each level, we adopt a selective classifier that can abstain from making a prediction at low confidence. Consequently, a positive (resp. negative) outcome leads to promotion (resp. demotion) of the agent to the next higher (resp. lower) level, while abstention keeps the agent at the same level. We fully characterize the agent's optimal instantaneous action under selective classifiers and compare the long-term properties and utility of the agent repeatedly following an optimal myopic policy of either no-improvement (never choose the improvement action) or no-gaming (never choose the gaming action). We further examine design principles over the sequence of classifiers that yield higher long-term utility for the latter policy, thereby effectively incentivizing genuine effort in the long run.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。