arXiv:2602.11944cs.LGcs.CY2026-02被引 2

提出量化模型预测分歧的指标,帮助AI系统符合欧盟法案对个体性能透明的要求。

Using predictive multiplicity to measure individual performance within the AI Act

  • 引入个体冲突率与δ-模糊性,量化单个案例中不同模型的预测差异
  • 发现高分歧个体在常见数据集上占比可达15%以上,影响决策可靠性
  • 为高风险AI系统提供可落地的评估规则,适合监管者与开发者使用

在构建决策支持型AI系统时,常出现预测多重性现象:不存在单一最优模型,多个相似整体准确率的模型在个别案例上的预测结果却可能完全不同。当决策直接影响个人时,这种不确定性极具问题。对于预测分歧大的个体,其结果可能因模型选择而改变,这违背了欧盟《人工智能法案》要求高风险AI系统不仅报告整体性能,还需报告特定个体表现的规定。本文旨在将预测多重性与法案中的准确性条款结合,提出具体实践建议:(1)分析法案对准确性的法律要求,指出纳入预测多重性信息有助于合规;(2)提出个体冲突比率和δ-模糊性作为量化个体层面模型分歧的工具;(3)基于计算可行性,设计简便可执行的评估规则;(4)建议将预测多重性信息向部署方公开,使其能判断特定个体输出是否可靠适用于实际场景。

原文摘要 · Abstract (English)

When building AI systems for decision support, one often encounters the phenomenon of predictive multiplicity: a single best model does not exist; instead, one can construct many models with similar overall accuracy that differ in their predictions for individual cases. Especially when decisions have a direct impact on humans, this can be highly unsatisfactory. For a person subject to high disagreement between models, one could as well have chosen a different model of similar overall accuracy that would have decided the person's case differently. We argue that this arbitrariness conflicts with the EU AI Act, which requires providers of high-risk AI systems to report performance not only at the dataset level but also for specific persons. The goal of this paper is to put predictive multiplicity in context with the EU AI Act's provisions on accuracy and to subsequently derive concrete suggestions on how to evaluate and report predictive multiplicity in practice. Specifically: (1) We introduce the AI Act's accuracy provisions and argue that incorporating information about predictive multiplicity could serve compliance with specific provisions for providers. (2) Based on this legally rigorous analysis, we suggest individual conflict ratios and $δ$-ambiguity as tools to quantify the disagreement between models on individual cases and to help detect individuals subject to conflicting predictions. (3) Based on computational insights, we derive easy-to-implement rules on how model providers could evaluate predictive multiplicity in practice. (4) Ultimately, we suggest that information about predictive multiplicity should be made available to deployers under the AI Act, enabling them to judge whether system outputs for specific individuals are reliable enough for their use case.

AI治理模型公平性法规合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。