即使无幻觉,医疗AI仍可能误导患者,因缺乏语境与关键信息。
Retrieval-augmented systems can be dangerous medical communicators
- 通过分析真实医疗查询,发现检索增强模型常忽略关键背景与来源。
- 患者从AI输出中理解的含义,往往与原文差异巨大,易产生误解。
- 建议引入沟通语用学和更深入的文档理解,提升可信度。
患者长期在线获取健康信息,如今越来越多依赖生成式AI回答医疗问题。鉴于医疗领域的高风险性,检索增强生成与引用定位等技术被广泛推广,以减少幻觉、提高回答准确性,并已集成至多数搜索引擎。本文指出,即便这些方法生成的内容在字面上准确且无虚构,仍可能极具误导性。患者从AI输出中获得的理解,往往与直接阅读原始文献或咨询专业医生的结果显著不同。通过对包括争议性诊断与手术安全性等主题的大规模查询分析,我们提供了量化与质性证据,表明现有系统存在严重缺陷:事实脱离语境、遗漏关键相关文献、强化患者误解或偏见。为此,我们提出一系列改进建议,如融入沟通语用学考量、增强对源文档的理解能力,这些思路亦可拓展至其他领域。
原文摘要 · Abstract (English)
Patients have long sought health information online, and increasingly, they are turning to generative AI to answer their health-related queries. Given the high stakes of the medical domain, techniques like retrieval-augmented generation and citation grounding have been widely promoted as methods to reduce hallucinations and improve the accuracy of AI-generated responses and have been widely adopted into search engines. This paper argues that even when these methods produce literally accurate content drawn from source documents sans hallucinations, they can still be highly misleading. Patients may derive significantly different interpretations from AI-generated outputs than they would from reading the original source material, let alone consulting a knowledgeable clinician. Through a large-scale query analysis on topics including disputed diagnoses and procedure safety, we support our argument with quantitative and qualitative evidence of the suboptimal answers resulting from current systems. In particular, we highlight how these models tend to decontextualize facts, omit critical relevant sources, and reinforce patient misconceptions or biases. We propose a series of recommendations -- such as the incorporation of communication pragmatics and enhanced comprehension of source documents -- that could help mitigate these issues and extend beyond the medical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。