arXiv:2504.08151cs.LG2025-04被引 1

通过自适应探索与中间动作,低成本缓解数据偏见问题。

Adaptive Bounded Exploration and Intermediate Actions for Data Debiasing

  • 用自适应边界控制探索范围,降低风险
  • 引入低成本中间标签,提升数据去偏效率
  • 适用于需要公平决策的高成本场景

算法决策性能高度依赖训练数据质量。数据中的偏见可能导致经济与伦理问题,引发对不同群体的不公对待。本文提出一种在分类任务中通过自适应且受限的探索来逐步去偏训练数据的算法,面对反馈成本高且存在截断的情况。所提方法在实现去偏目标(提升准确率与公平性)与探索风险之间取得平衡。具体而言,采用自适应边界限制探索区域,并利用低成本的中间动作获取噪声标签信息。理论上证明了该探索机制可在特定分布下帮助去偏数据,分析了算法公平干预措施与本方法的协同机制,并在合成数据与真实数据上通过数值实验验证了算法有效性。

原文摘要 · Abstract (English)

The performance of algorithmic decision rules is largely dependent on the quality of training datasets available to them. Biases in these datasets can raise economic and ethical concerns due to the resulting algorithms' disparate treatment of different groups. In this paper, we propose algorithms for sequentially debiasing the training dataset through adaptive and bounded exploration in a classification problem with costly and censored feedback. Our proposed algorithms balance between the ultimate goal of mitigating the impacts of data biases -- which will in turn lead to more accurate and fairer decisions, and the exploration risks incurred to achieve this goal. Specifically, we propose adaptive bounds to limit the region of exploration, and leverage intermediate actions which provide noisy label information at a lower cost. We analytically show that such exploration can help debias data in certain distributions, investigate how {algorithmic fairness interventions} can work in conjunction with our proposed algorithms, and validate the performance of these algorithms through numerical experiments on synthetic and real-world data.

数据去偏公平算法自适应探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。