arXiv:2502.15871cs.CYcs.AI2025-02EMNLP综述被引 41

系统梳理医疗大模型可信性问题,揭示关键风险与研究空白。

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

  • 从真实性、隐私、安全等六维度分析医疗大模型可信性挑战。
  • 总结现有评估框架,指出当前研究在公平性与可解释性上的不足。
  • 适合关注AI医疗落地的科研人员与临床决策者参考。

大语言模型(LLMs)在医疗领域应用前景广阔,可提升临床决策、医学研究和患者照护水平。然而,其在真实临床环境中的部署引发对可信性的关切,尤其涉及真实性、隐私、安全、鲁棒性、公平性和可解释性等维度。这些维度关乎模型输出的可靠性、无偏性与伦理合规性。尽管近期已有研究开发了评估基准与框架,但医疗大模型可信性仍缺乏系统性综述。本文通过全面回顾现有方法与解决方案,分析各维度对模型可靠性和伦理部署的影响,整合当前研究进展,并识别关键缺口。同时,针对多智能体协作、多模态推理及小型开源医学模型等新兴范式带来的新挑战进行探讨。旨在为未来研究提供方向,推动更可信、透明且具备临床可行性的医疗大模型发展。

原文摘要 · Abstract (English)

The application of large language models (LLMs) in healthcare holds significant promise for enhancing clinical decision-making, medical research, and patient care. However, their integration into real-world clinical settings raises critical concerns around trustworthiness, particularly around dimensions of truthfulness, privacy, safety, robustness, fairness, and explainability. These dimensions are essential for ensuring that LLMs generate reliable, unbiased, and ethically sound outputs. While researchers have recently begun developing benchmarks and evaluation frameworks to assess LLM trustworthiness, the trustworthiness of LLMs in healthcare remains underexplored, lacking a systematic review that provides a comprehensive understanding and future insights. This survey addresses that gap by providing a comprehensive review of current methodologies and solutions aimed at mitigating risks across key trust dimensions. We analyze how each dimension affects the reliability and ethical deployment of healthcare LLMs, synthesize ongoing research efforts, and identify critical gaps in existing approaches. We also identify emerging challenges posed by evolving paradigms, such as multi-agent collaboration, multi-modal reasoning, and the development of small open-source medical models. Our goal is to guide future research toward more trustworthy, transparent, and clinically viable LLMs.

大模型医疗AI可信性综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。