arXiv:2412.02000cs.LGcs.AI2024-12NeurIPS被引 7

通过因果分析识别恶意操纵模型的高风险用户

Who's Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation

  • 将每个用户的操纵程度建模为可量化参数,用因果推断方法排序
  • 证明了在不掌握用户目标函数的情况下仍可准确排名操纵程度
  • 在医保编码案例中成功识别出潜在操纵行为特征

在许多场景中,机器学习模型用于影响个体或实体的决策,这些实体可能通过操纵输入来获得更好结果,从而最大化自身利益。本文研究多主体环境下如何识别‘最严重操纵者’——即最积极进行策略性适应的个体。然而,在缺乏其效用函数信息的情况下,这一任务极具挑战。为此,我们提出一个框架,将每个代理的操纵倾向参数化为一个标量。我们证明该操纵参数仅部分可识别。通过将问题重新表述为因果效应估计问题(不同代理作为不同‘处理’),我们证明所有代理按操纵参数排序是可识别的。我们在合成数据上验证了因果推断在操纵检测中的有效性,并在一项美国诊断编码行为的案例研究中展示了本方法能有效突出与操纵相关的特征。

原文摘要 · Abstract (English)

In many settings, machine learning models may be used to inform decisions that impact individuals or entities who interact with the model. Such entities, or agents, may game model decisions by manipulating their inputs to the model to obtain better outcomes and maximize some utility. We consider a multi-agent setting where the goal is to identify the "worst offenders:" agents that are gaming most aggressively. However, identifying such agents is difficult without knowledge of their utility function. Thus, we introduce a framework in which each agent's tendency to game is parameterized via a scalar. We show that this gaming parameter is only partially identifiable. By recasting the problem as a causal effect estimation problem where different agents represent different "treatments," we prove that a ranking of all agents by their gaming parameters is identifiable. We present empirical results in a synthetic data study validating the usage of causal effect estimation for gaming detection and show in a case study of diagnosis coding behavior in the U.S. that our approach highlights features associated with gaming.

因果推断操纵检测多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。