arXiv:2412.01344cs.LGcs.GT2024-12被引 3

让机器学习模型适应策略性用户行为,提升预测稳定性。

Practical Performative Policy Learning with Strategic Agents

  • 将策略性行为建模为因果推断问题,不依赖强假设
  • 发现分布变化的低维结构,用可微分类器替代复杂映射
  • 仅需少量反馈数据即可高效优化,适合真实场景

本文研究可表演政策学习问题,即个体根据发布的政策调整自身特征以改善结果,导致数据分布内生变化。现有方法多依赖严格参数假设:战略分类中的微观效用模型或可表演预测中的宏观分布映射,严重限制可扩展性和泛化能力。我们将其视为复杂因果推断任务,放松对微观行为和宏观分布的参数假设。基于有限理性,揭示分布变化的低维结构,并构建从部署模型到分布偏移的因果路径中介。提出一种基于梯度的策略优化算法,以可微分类器替代高维分布映射。该算法高效利用批量反馈和有限操纵模式,在高维设置下实现比依赖赌博反馈或零阶优化的方法更高的样本效率。同时提供算法收敛的理论保证。大量且具有挑战性的实验验证了方法在实际应用中的有效性。

原文摘要 · Abstract (English)

This paper studies the performative policy learning problem, where agents adjust their features in response to a released policy to improve their potential outcomes, inducing an endogenous distribution shift. There has been growing interest in training machine learning models in strategic environments, including strategic classification and performative prediction. However, existing approaches often rely on restrictive parametric assumptions: micro-level utility models in strategic classification and macro-level data distribution maps in performative prediction, severely limiting scalability and generalizability. We approach this problem as a complex causal inference task, relaxing parametric assumptions on both micro-level agent behavior and macro-level data distribution. Leveraging bounded rationality, we uncover a practical low-dimensional structure in distribution shifts and construct an effective mediator in the causal path from the deployed model to the shifted data. We then propose a gradient-based policy optimization algorithm with a differentiable classifier as a substitute for the high-dimensional distribution map. Our algorithm efficiently utilizes batch feedback and limited manipulation patterns. Our approach achieves high sample efficiency compared to methods reliant on bandit feedback or zero-order optimization. We also provide theoretical guarantees for algorithmic convergence. Extensive and challenging experiments on high-dimensional settings demonstrate our method's practical efficacy.

策略学习因果推断可微优化分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。