系统评估医疗问答系统的可信度,覆盖六大核心维度。
Trustworthy Medical Question Answering: An Evaluation-Centric Survey
- 从事实性、鲁棒性等六方面构建可信度评估框架
- 对比主流评测基准,发现现有方法在多维评估上仍不足
- 适合关注AI医疗安全与可解释性的研究者参考
医疗问答系统中的可信度对患者安全、临床效果和用户信任至关重要。随着大语言模型(LLMs)在医疗场景中日益普及,其回答的可靠性直接影响临床决策与患者结局。然而,由于医疗数据复杂、临床场景敏感以及可信AI的多维度特性,实现全面可信度仍面临挑战。本文系统考察医疗QA中六个关键可信维度:事实性、鲁棒性、公平性、安全性、可解释性和校准性,综述了现有基于LLM的医疗QA系统如何评估这些维度。我们整理并比较了主要评测基准,并分析评估引导型改进技术,如检索增强的上下文生成、对抗性微调和安全对齐。最后,识别出可扩展专家评估、集成多维指标和真实世界部署研究等开放挑战,并提出未来研究方向,以推动LLM驱动的医疗QA在安全、可靠和透明方面的落地。
原文摘要 · Abstract (English)
Trustworthiness in healthcare question-answering (QA) systems is important for ensuring patient safety, clinical effectiveness, and user confidence. As large language models (LLMs) become increasingly integrated into medical settings, the reliability of their responses directly influences clinical decision-making and patient outcomes. However, achieving comprehensive trustworthiness in medical QA poses significant challenges due to the inherent complexity of healthcare data, the critical nature of clinical scenarios, and the multifaceted dimensions of trustworthy AI. In this survey, we systematically examine six key dimensions of trustworthiness in medical QA, i.e., Factuality, Robustness, Fairness, Safety, Explainability, and Calibration. We review how each dimension is evaluated in existing LLM-based medical QA systems. We compile and compare major benchmarks designed to assess these dimensions and analyze evaluation-guided techniques that drive model improvements, such as retrieval-augmented grounding, adversarial fine-tuning, and safety alignment. Finally, we identify open challenges-such as scalable expert evaluation, integrated multi-dimensional metrics, and real-world deployment studies-and propose future research directions to advance the safe, reliable, and transparent deployment of LLM-powered medical QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。