arXiv:2504.06791cs.LG2025-04被引 9

警惕AI解释的陷阱:好解释未必可信,设计不当反成风险

Beware of "Explanations" of AI

  • 从心理认知角度出发,强调解释需与用户心智模型对齐
  • 实证表明劣质解释可能引发误判、隐私泄露和信任倒退
  • 适合关注AI伦理、可解释性设计的研究者与政策制定者

理解日益复杂的AI系统所作决策与行为仍是关键挑战。这催生了可解释人工智能(XAI)研究,其被认为能增强信任、推动采纳并满足监管要求。然而,“好解释”的定义取决于目标、利益相关方与具体情境。尽管心理机制如心智模型对齐可提供指导,但实际应用受社会与技术因素制约。由于问题本身定义模糊,解释质量常不佳(如不忠实、无关或混乱),可能带来严重风险。不当解释不仅无法提升透明度,反而可能导致错误决策、隐私侵犯、操控甚至降低AI采用率。因此,我们警示各方:切勿盲目信赖AI解释——它们并非透明或负责任采纳的自动解药,滥用或忽视其局限性反而会加剧危害。关注这些警示,有助于未来研究提升解释的质量与实际影响。

原文摘要 · Abstract (English)

Understanding the decisions made and actions taken by increasingly complex AI system remains a key challenge. This has led to an expanding field of research in explainable artificial intelligence (XAI), highlighting the potential of explanations to enhance trust, support adoption, and meet regulatory standards. However, the question of what constitutes a "good" explanation is dependent on the goals, stakeholders, and context. At a high level, psychological insights such as the concept of mental model alignment can offer guidance, but success in practice is challenging due to social and technical factors. As a result of this ill-defined nature of the problem, explanations can be of poor quality (e.g. unfaithful, irrelevant, or incoherent), potentially leading to substantial risks. Instead of fostering trust and safety, poorly designed explanations can actually cause harm, including wrong decisions, privacy violations, manipulation, and even reduced AI adoption. Therefore, we caution stakeholders to beware of explanations of AI: while they can be vital, they are not automatically a remedy for transparency or responsible AI adoption, and their misuse or limitations can exacerbate harm. Attention to these caveats can help guide future research to improve the quality and impact of AI explanations.

可解释性AI伦理心智模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。