LLMs易受医学文献中的结果夸大影响,但可通过提示缓解。
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
- 测试22个LLM,发现其比人类更易受结果夸大影响
- 部分LLM在生成摘要时会隐式吸收并传播夸大内容
- 通过特定提示可有效降低夸大影响,提升判断准确性
医学研究在将新疗法转化为临床实践时面临诸多挑战。发表激励促使研究人员即使在证据模糊的情况下也倾向于呈现‘积极’结果,导致作者常在论文摘要中夸大研究结论。这种夸大可能影响临床医生对证据的理解,进而影响患者治疗决策。本文研究大型语言模型(LLMs)是否同样受此类夸大影响。我们评估了22个LLM,发现它们普遍比人类更易受夸大影响,并可能在生成的通俗摘要中隐式传播夸大内容。然而,我们也发现大多数LLM具备识别夸大能力,通过特定提示可有效减轻夸大对其输出的影响。
原文摘要 · Abstract (English)
Medical research faces well-documented challenges in translating novel treatments into clinical practice. Publishing incentives encourage researchers to present "positive" findings, even when empirical results are equivocal. Consequently, it is well-documented that authors often spin study results, especially in article abstracts. Such spin can influence clinician interpretation of evidence and may affect patient care decisions. In this study, we ask whether the interpretation of trial results offered by Large Language Models (LLMs) is similarly affected by spin. This is important since LLMs are increasingly being used to trawl through and synthesize published medical evidence. We evaluated 22 LLMs and found that they are across the board more susceptible to spin than humans. They might also propagate spin into their outputs: We find evidence, e.g., that LLMs implicitly incorporate spin into plain language summaries that they generate. We also find, however, that LLMs are generally capable of recognizing spin, and can be prompted in a way to mitigate spin's impact on LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。