arXiv:2504.16948cs.CYcs.AI2025-04被引 1

深度模型的复杂性存在内在解释障碍,难以通过现有方法突破。

Intrinsic Barriers to Explaining Deep Foundation Models

  • 从模型本质出发,分析解释困难的根源
  • 指出当前可解释性方法在深层模型面前力不从心
  • 提醒需重新思考模型验证与治理方式

深度基础模型(DFMs)展现出前所未有的能力,但其日益复杂的结构给理解其内部机制带来了巨大挑战,这关系到信任、安全与责任。面对解释难题,一个根本问题浮现:这些困难是暂时的技术瓶颈,还是由模型自身特性决定的内在障碍?本文通过剖析DFM的基本特征,审视当前可解释性方法在应对这一固有挑战时的局限性,探讨实现令人满意的解释是否可能,并反思如何重新构建对这些强大技术的验证与治理策略。

原文摘要 · Abstract (English)

Deep Foundation Models (DFMs) offer unprecedented capabilities but their increasing complexity presents profound challenges to understanding their internal workings-a critical need for ensuring trust, safety, and accountability. As we grapple with explaining these systems, a fundamental question emerges: Are the difficulties we face merely temporary hurdles, awaiting more sophisticated analytical techniques, or do they stem from \emph{intrinsic barriers} deeply rooted in the nature of these large-scale models themselves? This paper delves into this critical question by examining the fundamental characteristics of DFMs and scrutinizing the limitations encountered by current explainability methods when confronted with this inherent challenge. We probe the feasibility of achieving satisfactory explanations and consider the implications for how we must approach the verification and governance of these powerful technologies.

可解释性大模型可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。