arXiv:2502.04695cs.AIcs.CE2025-02中稿 · first EurIPS Works…被引 14

解释性评估需标准化,才能支撑AI可信监管与合规

Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance

  • 提出以度量为治理核心的XAI框架,增强可审计性
  • 强调解释性可防对齐造假,提升GPAI系统行为可信度
  • 适合关注AI合规、审计与监管的技术管理者

可靠的可解释性不仅是技术目标,更是私有AI治理的核心。随着AI进入高风险领域,审计机构、保险公司、认证组织和采购方需要标准化的评估指标来判断模型可信度。然而当前XAI评估指标碎片化且易被操纵,削弱了问责与合规性。我们主张标准化度量可作为治理基础,将可审计性和问责性嵌入AI系统以实现有效私有监督。基于已有XAI基准工作,我们识别出确保忠实性、抗篡改性和监管一致性的关键局限。可解释性可直接支持模型对齐,为通用目的AI(GPAI)系统提供可验证的行为完整性保障。这种解释性与对齐的关联,使XAI度量兼具技术与监管功能,有助于防止对齐造假这一日益受关注的问题。我们提出‘度量驱动治理’范式,将解释性评估作为私有AI治理的核心机制。该框架建立透明度、抗篡改性、可扩展性与法律一致性之间的分层联系,推动评估从模型内省延伸至系统性问责。通过概念整合与治理标准对接,我们勾勒出将解释性度量融入持续AI保障流程的路线图,服务于私有监督与监管双重需求。

原文摘要 · Abstract (English)

Reliable explainability is not only a technical goal but also a cornerstone of private AI governance. As AI models enter high-stakes sectors, private actors such as auditors, insurers, certification bodies, and procurement agencies require standardized evaluation metrics to assess trustworthiness. However, current XAI evaluation metrics remain fragmented and prone to manipulation, which undermines accountability and compliance. We argue that standardized metrics can function as governance primitives, embedding auditability and accountability within AI systems for effective private oversight. Building upon prior work in XAI benchmarking, we identify key limitations in ensuring faithfulness, tamper resistance, and regulatory alignment. Furthermore, interpretability can directly support model alignment by providing a verifiable means of ensuring behavioral integrity in General Purpose AI (GPAI) systems. This connection between interpretability and alignment positions XAI metrics as both technical and regulatory instruments that help prevent alignment faking, a growing concern among oversight bodies. We propose a Governance by Metrics paradigm that treats explainability evaluation as a central mechanism of private AI governance. Our framework introduces a hierarchical model linking transparency, tamper resistance, scalability, and legal alignment, extending evaluation from model introspection toward systemic accountability. Through conceptual synthesis and alignment with governance standards, we outline a roadmap for integrating explainability metrics into continuous AI assurance pipelines that serve both private oversight and regulatory needs.

可解释AIAI治理合规评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。