arXiv:2512.12500cs.HCcs.AI2025-12

XAI在皮肤科中对医生和公众影响不同,可能引发盲目依赖或偏见。

Explainable AI as a Double-Edged Sword in Dermatology: The Impact on Clinicians versus The Public

  • 用多模态大模型提供解释,对比不同群体反应。
  • 普通民众使用时,正确时准确率上升,错误时反而下降。
  • 医生不受解释影响,始终表现更稳定,适合临床部署。

人工智能正广泛渗透医疗领域,从医生助手到消费者应用。由于算法不透明影响人机交互,可解释人工智能(XAI)通过提供决策依据来缓解问题,但研究显示其可能引发过度依赖或偏见。我们基于两个大规模实验(623名普通公众;153名全科医生,PCPs),结合公平性诊断模型与不同XAI解释方式,考察了XAI尤其是多模态大语言模型(LLMs)对诊断性能的影响。结果显示,AI辅助在不同肤色患者间提升了准确率并减少了诊断差异。然而,LLM解释产生分化效应:公众在AI正确时准确率提升,在错误时反而下降;而经验丰富的医生则始终受益,不受AI准确性影响。先呈现AI建议也导致双方在AI出错时表现更差。这些发现揭示了XAI效果随用户专业水平和呈现时机变化,强调多模态大模型作为医疗AI中的‘双刃剑’特性,为未来人机协作系统设计提供依据。

原文摘要 · Abstract (English)

Artificial intelligence (AI) is increasingly permeating healthcare, from physician assistants to consumer applications. Since AI algorithm's opacity challenges human interaction, explainable AI (XAI) addresses this by providing AI decision-making insight, but evidence suggests XAI can paradoxically induce over-reliance or bias. We present results from two large-scale experiments (623 lay people; 153 primary care physicians, PCPs) combining a fairness-based diagnosis AI model and different XAI explanations to examine how XAI assistance, particularly multimodal large language models (LLMs), influences diagnostic performance. AI assistance balanced across skin tones improved accuracy and reduced diagnostic disparities. However, LLM explanations yielded divergent effects: lay users showed higher automation bias - accuracy boosted when AI was correct, reduced when AI erred - while experienced PCPs remained resilient, benefiting irrespective of AI accuracy. Presenting AI suggestions first also led to worse outcomes when the AI was incorrect for both groups. These findings highlight XAI's varying impact based on expertise and timing, underscoring LLMs as a "double-edged sword" in medical AI and informing future human-AI collaborative system design.

可解释AI皮肤科人机协作大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。