arXiv:2504.12424cs.HCcs.AI2025-04被引 5

让大模型当挑刺专家,揪出AI解释的漏洞

Don't Just Translate, Agitate: Using Large Language Models as Devil's Advocates for AI Explanations

  • 大模型不只翻译解释,更要主动质疑逻辑漏洞
  • 能暴露模型局限性,避免用户盲目信任
  • 适合需要深度理解AI决策的研究者与开发者

本文指出,当前可解释人工智能(XAI)研究中,大型语言模型(LLMs)常被用于将特征归因权重等技术输出转化为自然语言解释。尽管这提升了可读性,但研究表明,生成类人化解释未必提升用户理解,反而可能引发对AI系统的过度依赖。当LLMs在总结XAI结果时忽略模型局限、不确定性或不一致性时,会强化可解释性的假象,而非实现真正透明。本文主张:大模型不应仅作翻译,而应充当建设性挑刺者(即魔鬼代言人),主动提出替代解释、潜在偏见、训练数据限制,以及模型推理失效的案例。通过这种批判性角色,帮助用户更深入地审视AI决策,减少因误读或虚假解释导致的过度依赖。

原文摘要 · Abstract (English)

This position paper highlights a growing trend in Explainable AI (XAI) research where Large Language Models (LLMs) are used to translate outputs from explainability techniques, like feature-attribution weights, into a natural language explanation. While this approach may improve accessibility or readability for users, recent findings suggest that translating into human-like explanations does not necessarily enhance user understanding and may instead lead to overreliance on AI systems. When LLMs summarize XAI outputs without surfacing model limitations, uncertainties, or inconsistencies, they risk reinforcing the illusion of interpretability rather than fostering meaningful transparency. We argue that - instead of merely translating XAI outputs - LLMs should serve as constructive agitators, or devil's advocates, whose role is to actively interrogate AI explanations by presenting alternative interpretations, potential biases, training data limitations, and cases where the model's reasoning may break down. In this role, LLMs can facilitate users in engaging critically with AI systems and generated explanations, with the potential to reduce overreliance caused by misinterpreted or specious explanations.

可解释AI大模型批判性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。