arXiv:2511.05810cs.AIcs.CL2025-11

用混合模型提升疾病诊断可解释性,让AI报告像医生一样说人话。

DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis

  • 融合贝叶斯去卷积与eQTL先验,从基因数据推断细胞特异性表达
  • 阿尔茨海默病诊断准确率达88.0%,且输出可解释的临床报告
  • 用大模型生成面向医患的定制化诊断说明,增强可信度

构建可信的临床AI系统不仅需要高精度预测,还需透明、基于生物机制的解释。我们提出\texttt{DiagnoLLM},一种混合框架,整合贝叶斯去卷积、eQTL引导的深度学习和基于大模型的叙事生成,实现可解释的疾病诊断。该框架首先使用GP-unmix——一种基于高斯过程的分层模型——从批量和单细胞RNA测序数据中推断细胞类型特异的基因表达谱,并建模生物学不确定性。这些特征结合来自eQTL分析的调控先验信息,驱动神经分类器在阿尔茨海默病(AD)检测中达到88.0%的准确率。为支持人类理解与信任,我们引入一个基于大模型的推理模块,将模型输出转化为基于临床特征、归因信号和领域知识的受众定制化诊断报告。人工评估证实,这些报告准确、可操作,且对医生和患者均适切。研究结果表明,当大模型作为后处理推理器而非端到端预测器部署时,可在混合诊断流程中有效充当沟通桥梁。

原文摘要 · Abstract (English)

Building trustworthy clinical AI systems requires not only accurate predictions but also transparent, biologically grounded explanations. We present \texttt{DiagnoLLM}, a hybrid framework that integrates Bayesian deconvolution, eQTL-guided deep learning, and LLM-based narrative generation for interpretable disease diagnosis. DiagnoLLM begins with GP-unmix, a Gaussian Process-based hierarchical model that infers cell-type-specific gene expression profiles from bulk and single-cell RNA-seq data while modeling biological uncertainty. These features, combined with regulatory priors from eQTL analysis, power a neural classifier that achieves high predictive performance in Alzheimer's Disease (AD) detection (88.0\% accuracy). To support human understanding and trust, we introduce an LLM-based reasoning module that translates model outputs into audience-specific diagnostic reports, grounded in clinical features, attribution signals, and domain knowledge. Human evaluations confirm that these reports are accurate, actionable, and appropriately tailored for both physicians and patients. Our findings show that LLMs, when deployed as post-hoc reasoners rather than end-to-end predictors, can serve as effective communicators within hybrid diagnostic pipelines.

疾病诊断可解释AI大模型应用基因组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。