arXiv:2503.00269cs.LGcs.AI2025-03被引 13

用语义熵检测大模型医疗幻觉,尤其在妇产科更准更可靠。

Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy

  • 引入语义熵衡量内容意义层面的不一致性,识别幻觉。
  • 在英国妇产科考试数据集上,语义熵AUC达0.76,优于困惑度的0.62。
  • 临床专家验证显示其判别能力接近完美(AUC 0.97),适合资源有限场景。

大型语言模型(LLMs)在临床决策支持中潜力巨大,但其易产生虚假或误导性输出(即幻觉),阻碍了在医疗领域的广泛应用。在妇产科这类高风险领域,临床推理错误可能严重影响母婴结局,因此确保AI输出可靠性至关重要。传统不确定性度量如困惑度无法捕捉导致信息失真的语义层面不一致。本文评估了一种新型不确定性指标——语义熵(SE),用于检测AI生成的医学内容中的幻觉。基于源自英国皇家妇产科医师学会MRCOG考试的临床验证数据集,将SE与困惑度对比,结果显示SE表现更优,AUROC为0.76(95%置信区间:0.75–0.78),高于困惑度的0.62(0.60–0.65)。临床专家验证进一步证实其有效性,SE在不确定性判别上达到近完美的表现(AUROC: 0.97)。尽管语义聚类仅在30%的案例中成功,但语义熵仍可作为提升妇产科领域AI安全性的有力工具。研究表明,该方法有助于更可靠的AI临床整合,尤其适用于资源有限的环境,推动更安全高效的数字健康干预。

原文摘要 · Abstract (English)

Large language models (LLMs) hold substantial promise for clinical decision support. However, their widespread adoption in medicine, particularly in healthcare, is hindered by their propensity to generate false or misleading outputs, known as hallucinations. In high-stakes domains such as women's health (obstetrics & gynaecology), where errors in clinical reasoning can have profound consequences for maternal and neonatal outcomes, ensuring the reliability of AI-generated responses is critical. Traditional methods for quantifying uncertainty, such as perplexity, fail to capture meaning-level inconsistencies that lead to misinformation. Here, we evaluate semantic entropy (SE), a novel uncertainty metric that assesses meaning-level variation, to detect hallucinations in AI-generated medical content. Using a clinically validated dataset derived from UK RCOG MRCOG examinations, we compared SE with perplexity in identifying uncertain responses. SE demonstrated superior performance, achieving an AUROC of 0.76 (95% CI: 0.75-0.78), compared to 0.62 (0.60-0.65) for perplexity. Clinical expert validation further confirmed its effectiveness, with SE achieving near-perfect uncertainty discrimination (AUROC: 0.97). While semantic clustering was successful in only 30% of cases, SE remains a valuable tool for improving AI safety in women's health. These findings suggest that SE could enable more reliable AI integration into clinical practice, particularly in resource-limited settings where LLMs could augment care. This study highlights the potential of SE as a key safeguard in the responsible deployment of AI-driven tools in women's health, leading to safer and more effective digital health interventions.

语义熵大模型安全妇产科AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。