随机分类器能提升策略分类的准确率,且不会更差。
Strategic Classification with Randomised Classifiers
- 用随机分类器替代确定性分类器应对数据被策略性修改的问题。
- 在有限数据下,随机分类器的额外风险与确定性情况相似且可控制。
- 适合关注鲁棒性和实际部署中对抗策略行为的研究者。
我们研究策略分类问题,即学习者需基于被策略性修改的特征对个体进行分类。以往工作多局限于确定性分类器,本文首次理论分析允许使用随机分类器的情形。结果显示,在特定条件下,最优随机分类器的准确率优于最优确定性分类器,但绝不会更差。当有有限训练数据时,随机分类器上的策略经验风险最小化(Strategic ERM)的超额风险边界与确定性情形类似。在两种情况下,随着训练数据量增加,学习者得到的分类器风险均收敛至对应最优值,且收敛速率与独立同分布(i.i.d.)情形相同。我们的结论对比了此前相关理论工作,表明随机化可在不引入显著副作用的前提下缓解实践中可能遇到的部分问题。
原文摘要 · Abstract (English)
We consider the problem of strategic classification, where a learner must build a model to classify agents based on features that have been strategically modified. Previous work in this area has concentrated on the case when the learner is restricted to deterministic classifiers. In contrast, we perform a theoretical analysis of an extension to this setting that allows the learner to produce a randomised classifier. We show that, under certain conditions, the optimal randomised classifier can achieve better accuracy than the optimal deterministic classifier, but under no conditions can it be worse. When a finite set of training data is available, we show that the excess risk of Strategic Empirical Risk Minimisation over the class of randomised classifiers is bounded in a similar manner as the deterministic case. In both the deterministic and randomised cases, the risk of the classifier produced by the learner converges to that of the corresponding optimal classifier as the volume of available training data grows. Moreover, this convergence happens at the same rate as in the i.i.d. case. Our findings are compared with previous theoretical work analysing the problem of strategic classification. We conclude that randomisation has the potential to alleviate some issues that could be faced in practice without introducing any substantial downsides.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。