arXiv:2604.03254cs.CYcs.AI2026-04被引 1

AI准确率不是纯技术指标,而是依赖上下文的规范选择。

Is your AI Model Accurate Enough? The Difficult Choices Behind Rigorous AI Development and the EU AI Act

  • 准确率评估需在四类关键决策中权衡:选指标、平衡指标、用代表性数据、定阈值。
  • 技术实现中隐含对风险、错误和权衡的接受标准,影响实际合规性。
  • 为监管者、审计员和开发者提供将法律要求转化为技术实践的具体指引。

技术与法律讨论常将‘准确率’视为客观、可测量的纯技术属性。本文挑战这一观点,指出评估AI性能本质上依赖于情境相关的规范判断。这些技术-规范抉择对严谨的AI部署至关重要,决定了哪些错误优先处理、风险如何分配,以及如何解决竞争目标间的权衡。本文结合2024年欧盟《人工智能法案》中对高风险系统要求‘适当准确度’的规定,进行法律-技术分析,识别并解析四个核心决策环节:(1) 选择评估指标,(2) 平衡多个指标,(3) 在代表性数据上测量指标,(4) 确定接受阈值。针对每一项,研究其与法案中准确度要求及附带文档义务的关系,揭示其技术实现中嵌入的关于可接受风险、错误和权衡的隐含或明确假设,并通过案例和技术标准探讨其实际落地影响。通过显化准确率的技规范维度,本文推动更广泛的AI治理与监管跨学科讨论,并为监管者、审计员与开发者提供将(法律)安全要求转化为技术实践的具体指导。

原文摘要 · Abstract (English)

Technical and legal debates frequently suggest that "accuracy" is an objective, measurable, and purely technical property. We challenge this view, showing that evaluating AI performance fundamentally depends on context-dependent normative decisions. These techno-normative choices are crucial for rigorous AI deployment, as they determine which errors are prioritised, how risks are distributed, and how trade-offs between competing objectives are resolved. This paper provides a legal-technical analysis of the choices that shape how accuracy is defined, measured, and assessed, using the 2024 European Union AI Act -- which mandates an "appropriate level of accuracy" for high-risk systems -- as a primary case study. We identify and analyse four choices central to any robust performance evaluation: (1) selecting metrics, (2) balancing multiple metrics, (3) measuring metrics against representative data, and (4) determining acceptance thresholds. For each choice, we study its relationship to the AI Act's accuracy requirement and associated documentation obligations, show how its technical implementation embeds implicit or explicit assumptions about acceptable risks, errors, and trade-offs, and discuss the implications for the practical implementation of the AI Act by examples and related technical standards. By making the techno-normative dimensions of accuracy explicit, this paper contributes to broader interdisciplinary debates on AI governance and regulation, and offers specific guidance for regulators, auditors, and developers tasked with translating (legal) safety requirements into technical practice.

AI治理合规评估技术规范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。