arXiv:2503.16507cs.HCcs.AI2025-03被引 21

超99%的可解释AI论文未实证人类理解效果

Fewer Than 1% of Explainable AI Papers Validate Explainability with Humans

  • 分析1.8万篇论文,仅253篇提及人类参与评估
  • 真正开展人类实验的仅128篇,占比不足0.7%
  • 呼吁加强可解释AI研究中的人类验证环节

本研究对可解释人工智能(XAI)文献进行了大规模分析,评估其关于人类可解释性的声明。与专业图书馆员合作,识别出18,254篇包含可解释性、可理解性相关关键词的论文。其中仅有253篇提及人类参与评估XAI技术,而真正开展人类实验的仅128篇。这意味着在全部XAI论文中,提供人类可解释性实证证据的不足0.7%。研究揭示了可解释性宣称与实证验证之间的巨大差距,警示当前XAI研究的严谨性问题。作者呼吁加强人类评估,并公开文献检索方法以支持复现和进一步研究。

原文摘要 · Abstract (English)

This late-breaking work presents a large-scale analysis of explainable AI (XAI) literature to evaluate claims of human explainability. We collaborated with a professional librarian to identify 18,254 papers containing keywords related to explainability and interpretability. Of these, we find that only 253 papers included terms suggesting human involvement in evaluating an XAI technique, and just 128 of those conducted some form of a human study. In other words, fewer than 1% of XAI papers (0.7%) provide empirical evidence of human explainability when compared to the broader body of XAI literature. Our findings underscore a critical gap between claims of human explainability and evidence-based validation, raising concerns about the rigor of XAI research. We call for increased emphasis on human evaluations in XAI studies and provide our literature search methodology to enable both reproducibility and further investigation into this widespread issue.

可解释AI人类验证研究伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。