arXiv:2510.09705cs.LGcs.CY2025-10

用强化学习动态选特征,同时兼顾准确率和公平性。

A Multi-Component Reward Function with Policy Gradient for Automated Feature Selection with Dynamic Regularization and Bias Mitigation

  • 设计多组件奖励函数,融合预测性能与公平性
  • 代理自主选择特征子集,实时平衡准确与公平
  • 适合处理相关特征和隐蔽偏见的场景

静态特征剔除策略在隐藏依赖影响模型预测时往往无法有效防止偏见。为此,我们提出一种强化学习框架,将偏见缓解与自动化特征选择整合到单一学习过程中。与传统基于启发式的方法不同,我们的强化学习代理通过显式结合预测性能与公平性考量的奖励信号,自适应地选择特征。这种动态机制使模型在整个训练过程中平衡泛化能力、准确率与公平性,而非依赖预处理调整或事后校正。本文详细描述了多组件奖励函数的设计、代理动作空间在特征子集上的定义,以及该系统与集成学习的融合。旨在提供一种灵活且可泛化的特征选择方法,适用于预测因子相关且偏见可能隐性重现的环境。

原文摘要 · Abstract (English)

Static feature exclusion strategies often fail to prevent bias when hidden dependencies influence the model predictions. To address this issue, we explore a reinforcement learning (RL) framework that integrates bias mitigation and automated feature selection within a single learning process. Unlike traditional heuristic-driven filter or wrapper approaches, our RL agent adaptively selects features using a reward signal that explicitly integrates predictive performance with fairness considerations. This dynamic formulation allows the model to balance generalization, accuracy, and equity throughout the training process, rather than rely exclusively on pre-processing adjustments or post hoc correction mechanisms. In this paper, we describe the construction of a multi-component reward function, the specification of the agents action space over feature subsets, and the integration of this system with ensemble learning. We aim to provide a flexible and generalizable way to select features in environments where predictors are correlated and biases can inadvertently re-emerge.

特征选择强化学习公平性动态调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。