arXiv:2602.09326cs.LG2026-02被引 2

改进了数据与特征贡献度评估方法,能处理优先级依赖关系。

Priority-Aware Shapley Value

  • 引入硬约束和软权重的优先级机制,突破传统Shapley值假设。
  • 在MNIST/CIFAR10上实现更符合结构的数据估值,Census Income特征分配更合理。
  • 适合需要考虑数据重用或信任等级的场景,如医疗、金融建模。

Shapley值广泛用于模型无关的数据估值与特征归因,但其隐含假设是贡献者可互换。当贡献者存在依赖关系(如数据复用/增强或因果特征顺序)或需按信任、风险等调整贡献时,此假设可能失效。本文提出优先级感知的Shapley值(PASV),同时融入硬性优先级约束与贡献者特定的软权重。PASV适用于一般优先级结构,可退化为仅含优先级或仅含权重的特殊情形,且由自然公理唯一刻画。我们设计了高效的相邻交换Metropolis-Hastings采样器,支持大规模蒙特卡洛估计,并分析了极端权重下的渐近行为。在数据估值(MNIST/CIFAR10)和特征归因(Census Income)任务上的实验表明,PASV能实现更符合结构的分配结果,并通过提出的“优先级扫掠”实现实用的敏感性分析。

原文摘要 · Abstract (English)

Shapley values are widely used for model-agnostic data valuation and feature attribution, yet they implicitly assume contributors are interchangeable. This can be problematic when contributors are dependent (e.g., reused/augmented data or causal feature orderings) or when contributions should be adjusted by factors such as trust or risk. We propose Priority-Aware Shapley Value (PASV), which incorporates both hard precedence constraints and soft, contributor-specific priority weights. PASV is applicable to general precedence structures, recovers precedence-only and weight-only Shapley variants as special cases, and is uniquely characterized by natural axioms. We develop an efficient adjacent-swap Metropolis-Hastings sampler for scalable Monte Carlo estimation and analyze limiting regimes induced by extreme priority weights. Experiments on data valuation (MNIST/CIFAR10) and feature attribution (Census Income) demonstrate more structure-faithful allocations and a practical sensitivity analysis via our proposed "priority sweeping".

数据估值特征归因公平性机器学习解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。