arXiv:2511.00588cs.LGcs.AI2025-11被引 2

评估大模型在脊柱手术决策中的幻觉风险,提出多维度验证框架。

Diagnosing Hallucination Risk in AI Surgical Decision-Support: A Sequential Framework for Sequential Validation

  • 构建临床医生为中心的多维度评估框架,测试诊断、推荐、推理等能力。
  • 深求-R1表现最佳(总分86.03),但增强推理未必提升临床可靠性。
  • 发现模型输出流畅度与实际建议质量存在脱节,适合医疗AI安全评估者阅读。

大型语言模型(LLMs)在脊柱外科临床决策支持中具有变革潜力,但存在因幻觉导致的事实错误或语境错位,可能危及患者安全。本研究提出一种以临床医生为中心的框架,通过诊断精度、推荐质量、推理鲁棒性、输出连贯性和知识一致性五个维度量化幻觉风险。在30个专家验证的脊柱病例上评估了六种主流LLM。DeepSeek-R1整体表现最优(总分86.03 ± 2.08),尤其在创伤和感染等高风险领域表现突出。关键发现为:增强推理的模型版本并未普遍优于标准版本——Claude-3.7-Sonnet的扩展思维模式得分(80.79 ± 1.83)反而低于其标准版本(81.56 ± 1.92),表明单纯的链式思考不足以保证临床可靠性。多维度压力测试揭示模型特定弱点:在复杂度提升时,推荐质量下降7.4%;而理性(+2.0%)、可读性(+1.7%)和诊断准确率(+4.7%)略有提升,凸显输出流畅度与实际指导价值之间的严重偏离。研究建议将可解释性机制(如推理链可视化)融入临床流程,并建立面向安全的手术类LLM验证框架。

原文摘要 · Abstract (English)

Large language models (LLMs) offer transformative potential for clinical decision support in spine surgery but pose significant risks through hallucinations, which are factually inconsistent or contextually misaligned outputs that may compromise patient safety. This study introduces a clinician-centered framework to quantify hallucination risks by evaluating diagnostic precision, recommendation quality, reasoning robustness, output coherence, and knowledge alignment. We assessed six leading LLMs across 30 expert-validated spinal cases. DeepSeek-R1 demonstrated superior overall performance (total score: 86.03 $\pm$ 2.08), particularly in high-stakes domains such as trauma and infection. A critical finding reveals that reasoning-enhanced model variants did not uniformly outperform standard counterparts: Claude-3.7-Sonnet's extended thinking mode underperformed relative to its standard version (80.79 $\pm$ 1.83 vs. 81.56 $\pm$ 1.92), indicating extended chain-of-thought reasoning alone is insufficient for clinical reliability. Multidimensional stress-testing exposed model-specific vulnerabilities, with recommendation quality degrading by 7.4% under amplified complexity. This decline contrasted with marginal improvements in rationality (+2.0%), readability (+1.7%) and diagnosis (+4.7%), highlighting a concerning divergence between perceived coherence and actionable guidance. Our findings advocate integrating interpretability mechanisms (e.g., reasoning chain visualization) into clinical workflows and establish a safety-aware validation framework for surgical LLM deployment.

医疗AI幻觉检测大模型评估临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。