研究患者真实问诊问题中的错误假设,发现大模型难以识别日常医疗提问中的误导性信息。
What Patients Really Ask: Exploring the Effect of False Assumptions in Patient Information Seeking
- 从美国前200种处方药的‘相关搜索’中收集真实患者提问数据。
- 约三成问题包含错误假设或危险意图,且出现具有非随机规律。
- 现有大模型在真实问诊场景下识别错误假设能力显著不足,适合临床辅助工具开发者参考。
患者越来越多地使用大语言模型(LLMs)获取医疗相关信息。然而,当前针对问答任务的基准测试主要聚焦于医学考试题目,其风格和内容与患者真实提问存在显著差异。为弥合这一差距,我们通过查询美国前200种处方药,从Google的“人们也询问”功能中采集真实用户提问数据,构建了一个常见医疗问题数据集。所收集的问题中,相当一部分包含错误假设和潜在危险意图。我们证明,这些错误问题的出现并非完全随机,而是高度依赖于其历史问题中错误程度的累积。当前在其他基准上表现优异的大型语言模型,在识别日常医疗问题中的错误假设方面表现不佳。
原文摘要 · Abstract (English)
Patients are increasingly using large language models (LLMs) to seek answers to their healthcare-related questions. However, benchmarking efforts in LLMs for question answering often focus on medical exam questions, which differ significantly in style and content from the questions patients actually raise in real life. To bridge this gap, we sourced data from Google's People Also Ask feature by querying the top 200 prescribed medications in the United States, curating a dataset of medical questions people commonly ask. A considerable portion of the collected questions contains incorrect assumptions and dangerous intentions. We demonstrate that the emergence of these corrupted questions is not uniformly random and depends heavily on the degree of incorrectness in the history of questions that led to their appearance. Current LLMs that perform strongly on other benchmarks struggle to identify incorrect assumptions in everyday questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。