arXiv:2505.23954cs.LG2025-05

区分虚假申报与真实改变,量化代理在资源分配中的谎报程度。

Estimating Misreporting in the Presence of Genuine Modification: A Causal Perspective

  • 基于因果关系,通过对比操纵与未操纵数据集来识别谎报特征。
  • 在半合成和真实医保数据上验证,可准确估计谎报率。
  • 适用于资源分配场景中需防范策略性行为的系统设计者。

当机器学习模型用于资源配置时,受影响的个体可能为获得更好结果而策略性地改变自身特征。尽管已有研究关注策略性响应,但如何区分虚假申报与真实改变仍是根本挑战。本文提出一种因果驱动的方法,通过区分特征变化对下游变量(即因果后代)的影响,来识别并量化平均谎报程度。核心思想是:与真实改变不同,谎报特征不会对下游变量产生因果影响。我们通过比较操纵数据集与未操纵数据集下谎报特征对因果后代的因果效应,实现谎报率的可识别性,并刻画了估计器的方差。在半合成及真实医保数据集上进行实证验证,结果表明该方法可在现实场景中有效识别谎报行为。

原文摘要 · Abstract (English)

In settings where ML models are used to inform the allocation of resources, agents affected by the allocation decisions might have an incentive to strategically change their features to secure better outcomes. While prior work has studied strategic responses broadly, disentangling misreporting from genuine modification remains a fundamental challenge. In this paper, we propose a causally-motivated approach to identify and quantify how much an agent misreports on average by distinguishing deceptive changes in their features from genuine modification. Our key insight is that, unlike genuine modification, misreported features do not causally affect downstream variables (i.e., causal descendants). We exploit this asymmetry by comparing the causal effect of misreported features on their causal descendants as derived from manipulated datasets against those from unmanipulated datasets. We formally prove identifiability of the misreporting rate and characterize the variance of our estimator. We empirically validate our theoretical results using a semi-synthetic and real Medicare dataset with misreported data, demonstrating that our approach can be employed to identify misreporting in real-world scenarios.

因果推断策略性行为数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。