临床AI中,检索失败比幻觉更致命
Retrieval, not hallucinations, will be the limiting factor for LLM-based clinical AI tools

- 强调临床AI关键瓶颈是患者数据检索失败而非幻觉
- 指出检索错误影响诊断与治疗建议可靠性
- 适合临床医生与AI研发者关注系统性风险
大型语言模型(LLM)在临床人工智能中的错误讨论通常聚焦于精确性问题,如幻觉。本文从临床医生和AI研究者双重视角出发,呼吁将关注点转向召回率问题,特别是临床AI工具所需患者级数据的检索缺陷。文章梳理了各类检索错误类型及缓解策略,概述了大模型与检索技术的研究方向,并提供了检索评估的现状综述。
原文摘要 · Abstract (English)
Discussions around large language model (LLM) errors in clinical artificial intelligence (AI) generally center around precision errors like hallucinations. This perspective, targeting both clinicians and AI researchers, seeks to shift that discussion to recall errors, particularly in retrieval of patient-level data needed for many clinical AI tools. The perspective outlines types of errors and mitigation strategies, describes research directions in LLMs and retrieval, and provides an overview of retrieval evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。