arXiv:2512.18177cs.AIcs.CV2025-12中稿 · Asilomar Conferenc…

用知识引导的视觉模型提升医疗影像诊断的准确与可解释性

NEURO-GUARD: Neuro-Symbolic Generalization and Unbiased Adaptive Routing for Diagnostics -- Explainable Medical AI

  • 结合ViT与语言模型,通过临床知识自校验生成特征提取代码
  • 在4个数据集上比纯ViT提升6.2%准确率(84.69% vs. 78.4%)
  • 适合需要可解释性与跨域泛化的临床AI研发者

准确且可解释的基于图像的诊断仍是医学AI的核心挑战,尤其在数据有限、视觉线索微弱、临床决策高风险的场景中。现有视觉模型多依赖纯数据驱动学习,产生黑箱预测,可解释性差且跨域泛化能力弱,限制了实际临床应用。我们提出NEURO-GUARD,一种融合视觉变换器(ViTs)与语言驱动推理的知识引导视觉框架,以提升性能、透明度和领域鲁棒性。该框架采用检索增强生成(RAG)机制实现自我验证:大语言模型(LLM)迭代生成、评估并优化医学图像的特征提取代码。通过嵌入临床指南与专家知识,系统逐步提升特征检测与分类能力,超越纯数据驱动基线。在糖尿病视网膜病变分类任务中,于四个基准数据集APTOs、EyePACS、Messidor-1和Messidor-2上的实验表明,NEURO-GUARD相比纯ViT基线准确率提升6.2%(84.69% vs. 78.4%),跨域泛化能力提高5%。在基于MRI的癫痫检测任务中也验证了其跨域稳健性,持续优于现有方法。总体而言,NEURO-GUARD将符号化医学推理与非符号化视觉学习相融合,实现了可解释、知识感知且泛化能力强的医学图像诊断,在多个数据集上达到顶尖性能。

原文摘要 · Abstract (English)

Accurate yet interpretable image-based diagnosis remains a central challenge in medical AI, particularly in settings characterized by limited data, subtle visual cues, and high-stakes clinical decision-making. Most existing vision models rely on purely data-driven learning and produce black-box predictions with limited interpretability and poor cross-domain generalization, hindering their real-world clinical adoption. We present NEURO-GUARD, a novel knowledge-guided vision framework that integrates Vision Transformers (ViTs) with language-driven reasoning to improve performance, transparency, and domain robustness. NEURO-GUARD employs a retrieval-augmented generation (RAG) mechanism for self-verification, in which a large language model (LLM) iteratively generates, evaluates, and refines feature-extraction code for medical images. By grounding this process in clinical guidelines and expert knowledge, the framework progressively enhances feature detection and classification beyond purely data-driven baselines. Extensive experiments on diabetic retinopathy classification across four benchmark datasets APTOS, EyePACS, Messidor-1, and Messidor-2 demonstrate that NEURO-GUARD improves accuracy by 6.2% over a ViT-only baseline (84.69% vs. 78.4%) and achieves a 5% gain in domain generalization. Additional evaluations on MRI-based seizure detection further confirm its cross-domain robustness, consistently outperforming existing methods. Overall, NEURO-GUARD bridges symbolic medical reasoning with subsymbolic visual learning, enabling interpretable, knowledge-aware, and generalizable medical image diagnosis while achieving state-of-the-art performance across multiple datasets.

医疗AI可解释性知识融合视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。