arXiv:2509.08592cs.LGcs.AI2025-09中稿 · the first EurIPS W…被引 3

将可解释性作为设计原则,构建可信的AI治理基础设施。

Interpretability as Alignment: Making Internal Understanding a Design Principle

  • 把可解释性嵌入模型架构,实现审计可验证
  • 结合因果抽象与实证基准,提升模型透明度
  • 适合关注AI治理与问责的机构与开发者

前沿AI系统需要能够验证内部对齐性的治理机制,而不仅仅是行为合规。私人治理机制如审计、认证、保险和采购正在补充公共监管,但这些机制需要能生成关于模型行为的可验证因果证据的技术基础。本文主张,机制性可解释性提供了这一基础。我们不将可解释性视为事后解释,而是将其作为设计约束,将可审计性、来源追溯性和有限透明性嵌入模型架构中。结合因果抽象理论与实证基准MIB和LoBOX,我们阐述了以可解释性为先的模型如何支撑私人保证流程和角色校准的透明框架。这一重构使可解释性成为连接技术可靠性与制度问责的私有AI治理基础设施。

原文摘要 · Abstract (English)

Frontier AI systems require governance mechanisms that can verify internal alignment, not just behavioral compliance. Private governance mechanisms audits, certification, insurance, and procurement are emerging to complement public regulation, but they require technical substrates that generate verifiable causal evidence about model behavior. This paper argues that mechanistic interpretability provides this substrate. We frame interpretability not as post-hoc explanation but as a design constraint embedding auditability, provenance, and bounded transparency within model architectures. Integrating causal abstraction theory and empirical benchmarks such as MIB and LoBOX, we outline how interpretability-first models can underpin private assurance pipelines and role-calibrated transparency frameworks. This reframing situates interpretability as infrastructure for private AI governance bridging the gap between technical reliability and institutional accountability.

可解释性AI治理因果抽象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。