同一医学证据下,提问方式不同会导致大模型回答矛盾,需重视提示鲁棒性。
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
- 在相同临床试验摘要基础上,对比正反向提问与技术/通俗语言的差异
- 正反提问导致结论矛盾的概率显著高于同向提问,多轮对话中矛盾更严重
- 提示鲁棒性应成为医疗问答系统评估的关键标准
患者越来越多地使用大型语言模型(LLMs)提出复杂且难以清晰表达的医学问题。然而,LLMs对提示语的表述方式敏感,可能受提问方式影响。理想情况下,即使提示语不同,只要基于相同证据,其回答也应保持一致。本研究在控制的检索增强生成(RAG)环境下,通过专家精选的文档进行医学问答评估,考察患者提问的两个维度:问题框架(正向与负向)和语言风格(专业与通俗)。我们构建了6,614对基于临床试验摘要的提问对,并评估8个LLM的响应一致性。结果表明,正向与负向提问对产生矛盾结论的概率显著高于同框架提问对。该框架效应在多轮对话中进一步加剧,持续说服导致不一致上升。语言风格与框架之间无显著交互作用。研究证明,仅通过改变提问方式,就能系统性影响基于相同证据的LLM医疗问答结果,强调了在高风险场景中提示鲁棒性作为RAG系统评估标准的重要性。
原文摘要 · Abstract (English)
Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt phrasings and can be influenced by the way questions are worded. Ideally, LLMs should respond consistently regardless of phrasing, particularly when grounded in the same underlying evidence. We investigate this through a systematic evaluation in a controlled retrieval-augmented generation (RAG) setting for medical question answering (QA), where expert-selected documents are used rather than retrieved automatically. We examine two dimensions of patient query variation: question framing (positive vs. negative) and language style (technical vs. plain language). We construct a dataset of 6,614 query pairs grounded in clinical trial abstracts and evaluate response consistency across eight LLMs. Our findings show that positively- and negatively-framed pairs are significantly more likely to produce contradictory conclusions than same-framing pairs. This framing effect is further amplified in multi-turn conversations, where sustained persuasion increases inconsistency. We find no significant interaction between framing and language style. Our results demonstrate that LLM responses in medical QA can be systematically influenced through query phrasing alone, even when grounded in the same evidence, highlighting the importance of phrasing robustness as an evaluation criterion for RAG-based systems in high-stakes settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。