arXiv:2507.01282cs.AIcs.HC2025-07被引 2

混合可解释系统让AI更懂临床,提升医生信任与诊疗效果。

Beyond Black-Box AI: Interpretable Hybrid Systems for Dementia Care

  • 结合统计学习与专家规则的混合模型增强可解释性
  • 实证显示此类系统更契合临床流程,如PEIRS、ATHENA-CDS
  • 未来应以医生理解度和患者结果衡量AI成功与否

大型语言模型(LLMs)虽在基准测试中表现亮眼,但在痴呆症诊断与护理的实际临床场景中尚未带来可测量的改善。独立机器学习模型擅长模式识别却难提供可操作、可解释的建议,削弱医生信任;医生使用LLMs也未提升诊断准确率或速度。主要瓶颈在于数据驱动范式:黑箱输出缺乏透明性,易产生幻觉,且因果推理能力弱。将统计学习与专家规则相结合,并全程纳入临床医生参与的混合方法,有助于恢复可解释性,更贴合现有临床工作流,如PEIRS和ATHENA-CDS系统所示。未来决策支持应优先实现预测与临床因果之间的解释性连结,通过神经符号或混合AI融合大模型语言能力与人类因果知识。尽管可解释AI与神经符号AI已成研究方向,仍依赖数据驱动整合,而非人机协同。未来研究需不仅关注准确率,更要评估医生理解度、工作流适配度及患者结局改善。提升人机交互有效性是推动AI进入临床实践的关键。

原文摘要 · Abstract (English)

The recent boom of large language models (LLMs) has re-ignited the hope that artificial intelligence (AI) systems could aid medical diagnosis. Yet despite dazzling benchmark scores, LLM assistants have yet to deliver measurable improvements at the bedside. This scoping review aims to highlight the areas where AI is limited to make practical contributions in the clinical setting, specifically in dementia diagnosis and care. Standalone machine-learning models excel at pattern recognition but seldom provide actionable, interpretable guidance, eroding clinician trust. Adjacent use of LLMs by physicians did not result in better diagnostic accuracy or speed. Key limitations trace to the data-driven paradigm: black-box outputs which lack transparency, vulnerability to hallucinations, and weak causal reasoning. Hybrid approaches that combine statistical learning with expert rule-based knowledge, and involve clinicians throughout the process help bring back interpretability. They also fit better with existing clinical workflows, as seen in examples like PEIRS and ATHENA-CDS. Future decision-support should prioritise explanatory coherence by linking predictions to clinically meaningful causes. This can be done through neuro-symbolic or hybrid AI that combines the language ability of LLMs with human causal expertise. AI researchers have addressed this direction, with explainable AI and neuro-symbolic AI being the next logical steps in further advancement in AI. However, they are still based on data-driven knowledge integration instead of human-in-the-loop approaches. Future research should measure success not only by accuracy but by improvements in clinician understanding, workflow fit, and patient outcomes. A better understanding of what helps improve human-computer interactions is greatly needed for AI systems to become part of clinical practice.

可解释AI痴呆诊断混合系统临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。