arXiv:2503.22746cs.CLcs.AI2025-03被引 11

用户提问方式会影响医疗大模型诊断准确率,尤其权威语气误导危害最大。

Susceptibility of Large Language Models to User-Driven Factors in Medical Queries

  • 测试不同提问方式对模型输出的影响,包括误导信息和缺失临床数据。
  • 权威语气使模型准确率下降最明显,缺少检验结果时性能骤降。
  • 建议用户使用规范提示词并提供完整病历,尤其复杂病例需谨慎提问。

大型语言模型在医疗领域应用日益广泛,但其可靠性受用户提问方式影响显著,如问题表述、临床信息完整性等。本研究通过两个实验评估误导性信息呈现方式、来源权威性、模型角色设定及关键临床信息缺失对诊断准确性的影响。实验一引入不同肯定程度的外部错误信息(扰动测试),实验二移除特定患者信息类别(消融测试)。基于MedQA与Medbullets公开数据集,评估了GPT-4o、Claude 3.5 Sonnet、Claude 3.5 Haiku、Gemini 1.5 Pro、Gemini 1.5 Flash、LLaMA 3 8B、LLaMA 3 Med42 8B、DeepSeek R1 8B等模型。所有模型均易受用户引导影响,其中专有模型对明确且权威的表述尤为敏感;具有强断言语气的信息对准确率影响最大。在消融测试中,缺失体格检查和实验室结果导致性能下降最严重。尽管专有模型基础准确率更高,但在误导信息下表现急剧下滑。研究强调需使用结构化提示词并提供完整临床背景,用户应避免以权威口吻传递错误信息,并确保提供全面病史。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in healthcare, but their reliability is heavily influenced by user-driven factors such as question phrasing and the completeness of clinical information. In this study, we examined how misinformation framing, source authority, model persona, and omission of key clinical details affect the diagnostic accuracy and reliability of LLM outputs. We conducted two experiments: one introducing misleading external opinions with varying assertiveness (perturbation test), and another removing specific categories of patient information (ablation test). Using public datasets (MedQA and Medbullets), we evaluated proprietary models (GPT-4o, Claude 3.5 Sonnet, Claude 3.5 Haiku, Gemini 1.5 Pro, Gemini 1.5 Flash) and open-source models (LLaMA 3 8B, LLaMA 3 Med42 8B, DeepSeek R1 8B). All models were vulnerable to user-driven misinformation, with proprietary models especially affected by definitive and authoritative language. Assertive tone had the greatest negative impact on accuracy. In the ablation test, omitting physical exam findings and lab results caused the most significant performance drop. Although proprietary models had higher baseline accuracy, their performance declined sharply under misinformation. These results highlight the need for well-structured prompts and complete clinical context. Users should avoid authoritative framing of misinformation and provide full clinical details, especially for complex cases.

大模型医疗问答提示工程可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。