arXiv:2601.02123cs.CLcs.AI2026-01

让大模型回答更贴合患者实际需求,提升临床实用性。

DeCode: Decoupling Content and Delivery for Medical QA

  • 分离内容与表达方式,使回答更契合患者个体情况。
  • 零样本性能从28.4%提升至49.8%,超越现有方法。
  • 无需训练、适配任何大模型,适合医疗问答场景。

大型语言模型(LLMs)具备强大的医学知识,能生成事实准确的回答。然而,现有模型常忽视个体患者背景,导致回答虽临床正确却难以满足实际需求。本文提出 DeCode(Decoupling Content and Delivery),一种无需训练、适用于任意模型的框架,可使现有 LLM 在临床场景中生成更具情境适配性的回答。我们在 OpenAI HealthBench 上评估该方法,这是一个综合性强、挑战性高的基准,用于检验 LLM 回答的临床相关性和有效性。DeCode 将零样本性能从 28.4% 提升至 49.8%,达到当前最优水平。实验表明,该方法显著提升了 LLM 在临床问答中的表现。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit strong medical knowledge and can generate factually accurate responses. However, existing models often fail to account for individual patient contexts, producing answers that are clinically correct yet poorly aligned with patients' needs. In this work, we introduce DeCode (Decoupling Content and Delivery), a training-free, model-agnostic framework that adapts existing LLMs to produce contextualized answers in clinical settings. We evaluate DeCode on OpenAI HealthBench, a comprehensive and challenging benchmark designed to assess clinical relevance and validity of LLM responses. DeCode boosts zero-shot performance from 28.4% to 49.8% and achieves new state-of-the-art compared to existing methods. Experimental results suggest the effectiveness of DeCode in improving clinical question answering of LLMs.

医疗问答大模型情境适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。