arXiv:2508.00285cs.CL2025-08

让大模型看病更靠谱,通过病因注意力监督提升诊断可信度

Enhancing Trustworthy Clinical Diagnosis Decision-Making in Large Language Models via Etiology-Aware Attention Supervision

  • 用临床指南构建病因结构,监督模型关注关键诊断证据
  • 诊断准确率提升15.65%,注意力更集中于病因相关信息
  • 适合医疗AI可信性研究者和临床决策系统开发者

大型语言模型在医学文本理解与生成方面表现强劲,但在以诊断为导向的任务中,其可信度受限于缺乏对临床诊断证据内部注意力机制的结构化引导。本文提出病因感知注意力监督框架,基于权威临床指南构建三种急腹症(急性阑尾炎、急性胰腺炎、急性胆囊炎)的临床病因谱(CES),并据此识别与病因一致的注意力头。在此基础上,设计结构引导的参数高效微调方法,通过额外监督损失引导注意力分布聚焦于临床相关证据,不修改基础模型架构。在一致性诊断队列上的实验显示,该框架使平均诊断准确率提升15.65%。基于注意力的指标(推理聚焦得分、推理注意力频率)表明模型更集中关注病因相关证据。在存在临床分歧的离群队列上,外部评估进一步验证了诊断性能提升的鲁棒性。

原文摘要 · Abstract (English)

Objective: Large Language Models (LLMs) have demonstrated strong capabilities in medical text understanding and generation. However, their trustworthiness in diagnosis-oriented medical tasks remains constrained by the lack of structured guidance on how clinically relevant diagnostic evidence is internally attended to and utilized during model learning. Method: We propose an Etiology-Aware Attention Supervision framework that introduces structured etiological information as an external supervisory signal for training large language models. Specifically, we construct Clinical Etiology Schema (CES) derived from authoritative clinical guidelines for three acute abdominal conditions: acute appendicitis, acute pancreatitis, and acute cholecystitis. Based on CES annotations, we develop an Etiology-Aware Head Identification strategy to identify attention heads that consistently align with etiological evidence. Building on this analysis, we design a structure-guided parameter-efficient fine-tuning approach that steers attention distributions toward clinically relevant evidence through an additional supervision loss, without modifying the base model architecture. Result: Experiments conducted on a Consistent Diagnosis Cohort demonstrate that the proposed framework improves average diagnostic accuracy by 15.65% compared with baseline models. Attention-based metrics, including Inference Focus Score and Inference Attention Frequency, show more concentrated attention on etiologically relevant evidence. External evaluation on a Discrepant Diagnosis Cohort further confirms the robustness of diagnostic performance improvements under real-world clinical inconsistencies.

医疗AI注意力监督诊断可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。