arXiv:2604.05348cs.AI2026-04

用眼底图像证据检测医疗大模型幻觉,提升诊断安全性。

From Retinal Evidence to Safe Decisions: RETINA-SAFE and ECRT for Hallucination Risk Triage in Medical LLMs

  • 构建眼底病证据基准RETINA-SAFE,分三类证据情境测试幻觉。
  • 提出ECRT框架,两阶段识别风险,准确率比基线高0.15~0.19。
  • 适合关注医疗AI可解释性与安全性的研究者和开发者。

医疗大语言模型的幻觉问题在证据不足或矛盾时尤为严重。本文以糖尿病视网膜病变决策场景为例,提出与眼底分级记录对齐的证据基准RETINA-SAFE,包含12,522个样本,分为三类证据关系任务:E-Align(证据一致)、E-Conflict(证据冲突)和E-Gap(证据不足)。进一步提出ECRT(Evidence-Conditioned Risk Triage)——一种两阶段白盒检测框架:第一阶段进行安全/不安全风险分诊,第二阶段将不安全案例细分为矛盾驱动与证据缺失两类。ECRT利用上下文/无上下文条件下的内部表示与逻辑值偏移,采用类别平衡训练。在多个骨干模型上,基于证据分组(非患者独立)划分,ECRT在第一阶段风险分诊中表现优异,相比外部不确定性与自一致性基线提升0.15至0.19的平衡准确率,较最强适应型监督基线提升0.02至0.07,并持续优于单阶段白盒消融实验。结果表明,基于眼底证据的白盒内部信号是实现可解释医疗大模型风险分诊的可行路径。

原文摘要 · Abstract (English)

Hallucinations in medical large language models (LLMs) remain a safety-critical issue, particularly when available evidence is insufficient or conflicting. We study this problem in diabetic retinopathy (DR) decision settings and introduce RETINA-SAFE, an evidence-grounded benchmark aligned with retinal grading records, comprising 12,522 samples. RETINA-SAFE is organized into three evidence-relation tasks: E-Align (evidence-consistent), E-Conflict (evidence-conflicting), and E-Gap (evidence-insufficient). We further propose ECRT (Evidence-Conditioned Risk Triage), a two-stage white-box detection framework: Stage 1 performs Safe/Unsafe risk triage, and Stage 2 refines unsafe cases into contradiction-driven versus evidence-gap risks. ECRT leverages internal representation and logit shifts under CTX/NOCTX conditions, with class-balanced training for robust learning. Under evidence-grouped (not patient-disjoint) splits across multiple backbones, ECRT provides strong Stage-1 risk triage and explicit subtype attribution, improves Stage-1 balanced accuracy by +0.15 to +0.19 over external uncertainty and self-consistency baselines and by +0.02 to +0.07 over the strongest adapted supervised baseline, and consistently exceeds a single-stage white-box ablation on Stage-1 balanced accuracy. These findings support white-box internal signals grounded in retinal evidence as a practical route to interpretable medical LLM risk triage.

医疗LLM幻觉检测可解释性眼底病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。