arXiv:2410.20636cs.CRcs.AI2024-10被引 1

LLM可作医疗决策的第二意见,生成全面鉴别诊断。

Language Models And A Second Opinion Use Case: The Pocket Professional

  • 用183个复杂病例测试多款LLM,对比医生群体共识。
  • 基础模型整体准确率超80%,但复杂病例仅43%准确。
  • 适合辅助医生防认知偏差,非替代主诊工具。

本研究检验大型语言模型(LLMs)作为专业决策中正式第二意见工具的作用,聚焦于即使经验丰富的医生也需同行会诊的复杂医疗案例。分析了20个月间来自Medscape的183个挑战性病例,对比多个LLM与众包医师响应的表现。关键发现是最新基础模型在整体上表现优异(相比共识意见准确率超过80%),超越多数人类在相同临床案例中的表现(涵盖450页患者病历与检测结果)。研究显示,LLM在简单病例中准确率高于81%,但在复杂场景中降至43%,尤其在引发医生广泛争议的案例中表现更差。结果表明,LLM更适合生成全面的鉴别诊断,而非作为主要诊断工具,有助于对抗认知偏见、减轻认知负荷,从而减少医疗错误。另纳入21个最高法院案例的对比法律数据集,虽其对LLM而言较易处理,但仍为人工智能提供第二意见提供了额外实证背景。研究还构建了一个新颖基准,供他人评估高争议问答在LLM与意见分歧的人类从业者间的可靠性。这些结果提示,LLM在专业场景中的最优部署方式可能与当前强调自动化常规任务的做法大相径庭。

原文摘要 · Abstract (English)

This research tests the role of Large Language Models (LLMs) as formal second opinion tools in professional decision-making, particularly focusing on complex medical cases where even experienced physicians seek peer consultation. The work analyzed 183 challenging medical cases from Medscape over a 20-month period, testing multiple LLMs' performance against crowd-sourced physician responses. A key finding was the high overall score possible in the latest foundational models (>80% accuracy compared to consensus opinion), which exceeds most human metrics reported on the same clinical cases (450 pages of patient profiles, test results). The study rates the LLMs' performance disparity between straightforward cases (>81% accuracy) and complex scenarios (43% accuracy), particularly in these cases generating substantial debate among human physicians. The research demonstrates that LLMs may be valuable as generators of comprehensive differential diagnoses rather than as primary diagnostic tools, potentially helping to counter cognitive biases in clinical decision-making, reduce cognitive loads, and thus remove some sources of medical error. The inclusion of a second comparative legal dataset (Supreme Court cases, N=21) provides added empirical context to the AI use to foster second opinions, though these legal challenges proved considerably easier for LLMs to analyze. In addition to the original contributions of empirical evidence for LLM accuracy, the research aggregated a novel benchmark for others to score highly contested question and answer reliability between both LLMs and disagreeing human practitioners. These results suggest that the optimal deployment of LLMs in professional settings may differ substantially from current approaches that emphasize automation of routine tasks.

医疗AI第二意见大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。