arXiv:2507.21188cs.LGcs.AI2025-07被引 2

发现临床大模型在细微输入变化下诊断易失稳,提出新评估框架检测潜在风险。

Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs

  • 通过结构化扰动模拟临床常见误判场景,系统测试模型隐空间鲁棒性。
  • 在真实临床数据上发现微小修改即引发诊断翻转,隐空间脆弱性普遍存在。
  • 适合医疗AI安全审计、模型开发者及临床决策系统验证者使用。

用于临床决策支持的大语言模型在面对小幅但具临床意义的输入变化(如症状遮蔽或发现否定)时常出现推理失败,尽管其在静态基准上表现良好。此类推理失误通常无法被标准NLP指标捕捉,因这些指标对驱动诊断不稳定的隐表示变化不敏感。本文提出几何感知评估框架LAPD(Latent Agentic Perturbation Diagnostics),系统探测临床大模型在结构化对抗编辑下的隐空间鲁棒性。在此框架中,引入模型无关的诊断翻转率(LDFR),量化嵌入在主成分分析降维后的隐空间中跨越决策边界时的表征不稳定性。临床文本通过基于诊断推理的结构化提示流程生成,并沿四种轴向扰动:遮蔽、否定、同义替换与数值变化,以模拟常见模糊与遗漏。我们在基础模型与临床专用模型上计算LDFR,发现即使表面变化极小,也存在显著的隐空间脆弱性。最后,在90条来自DiReCT基准(MIMIC-IV)的真实临床笔记上验证结果,确认LDFR在真实场景中的泛化能力。研究揭示了表面鲁棒性与语义稳定性之间的持续差距,强调在高风险临床AI中开展几何感知审计的重要性。

原文摘要 · Abstract (English)

LLMs for clinical decision support often fail under small but clinically meaningful input shifts such as masking a symptom or negating a finding, despite high performance on static benchmarks. These reasoning failures frequently go undetected by standard NLP metrics, which are insensitive to latent representation shifts that drive diagnosis instability. We propose a geometry-aware evaluation framework, LAPD (Latent Agentic Perturbation Diagnostics), which systematically probes the latent robustness of clinical LLMs under structured adversarial edits. Within this framework, we introduce Latent Diagnosis Flip Rate (LDFR), a model-agnostic diagnostic signal that captures representational instability when embeddings cross decision boundaries in PCA-reduced latent space. Clinical notes are generated using a structured prompting pipeline grounded in diagnostic reasoning, then perturbed along four axes: masking, negation, synonym replacement, and numeric variation to simulate common ambiguities and omissions. We compute LDFR across both foundation and clinical LLMs, finding that latent fragility emerges even under minimal surface-level changes. Finally, we validate our findings on 90 real clinical notes from the DiReCT benchmark (MIMIC-IV), confirming the generalizability of LDFR beyond synthetic settings. Our results reveal a persistent gap between surface robustness and semantic stability, underscoring the importance of geometry-aware auditing in safety-critical clinical AI.

临床大模型隐空间脆弱性诊断稳定性安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。