NEURON让临床AI模型解释更可信,把数据变成医生能懂的自然语言。
NEURON: A Neuro-symbolic System for Grounded Clinical Explainability

- 融合医学术语库与机器学习,用结构化知识增强模型理解力
- 在心衰死亡预测上AUC提升至0.74-0.77,比传统方法高0.029
- 生成的解释更符合医生认知,人类评估得分达0.807
临床AI应用受限于高性能模型的黑箱特性,缺乏专业级可解释性所需的本体基础与叙事透明度。我们提出NEURON,一种神经符号系统,旨在提升预测可靠性与临床可解释性。该系统将SNOMED CT本体引导的结构表示与机器学习模型结合,弥合原始数据与医学术语之间的鸿沟。为实现人机对齐交互,系统采用基于检索增强生成(RAG)的大型语言模型(LLM)层,将SHAP特征归因与患者特异性临床笔记整合为连贯、自然语言的解释。在MIMIC-IV数据集上针对急性心力衰竭死亡率预测进行验证,NEURON将AUC从0.71–0.74提升至0.74–0.77,并在人类对齐指标上显著优于基于SHAP的解释(0.807 vs. 0.558)。结果表明,NEURON为部署可信赖、以人为本的智能医疗应用提供了稳健且可扩展的工程方案。
原文摘要 · Abstract (English)
Clinical AI adoption is hindered by the black-box/grey-box nature of high-performing models, which lack the ontological grounding and narrative transparency required for professional-level explainability. We present NEURON, a neuro-symbolic system designed to enhance both predictive reliability and clinical interpretability. NEURON integrates SNOMED CT ontology-informed structural representations with machine learning models to bridge the gap between raw data and medical nomenclature. To facilitate human-aligned interaction, the system utilizes a Retrieval-Augmented Generation (RAG) grounded Large Language Model (LLM) layer to synthesize SHAP feature attributions and patient-specific clinical notes into coherent, natural-language explanations. Validated on the MIMIC-IV dataset for Acute Heart Failure mortality prediction, NEURON improved the AUC from 0.71-0.74 to 0.74-0.77 and substantially outperformed SHAP-based explanations in human-aligned metrics (0.807 vs. 0.558). Our results demonstrate that NEURON offers a robust, scalable engineering solution for deploying trustworthy, human-centered connected health applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。