医疗大模型评估应关注初始咨询行为,而非仅看最终诊断。
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
- 在真实咨询初期就测试模型表现,而非等病情明确后
- 无引导时9/12案例提前给自疗建议,有引导则0/12
- 适合关注临床交互质量的研究者与开发者
医疗大语言模型常在临床问题已明确后才被评估,但真实咨询往往始于模糊、简化或错误表述的问题。我们对三个API模型在四份医师编写的多轮病例情景中,分别在基础条件和进入照护指令条件下进行测试,生成24份固定脚本对话;另两例采用自适应标准化患者模拟,生成12份对话。在无指令条件下,9/12案例在患者回答前就给出自我护理建议;有指令时则为0/12。结构化交接报告在无指令时未出现(0/12),有指令时则出现在10/12案例中。指令改变了对话顺序与记录方式,但未能稳定获取关键信息。因此,应直接通过首次接触行为来评估‘预构型差距’,而非依赖诊断准确率或最终答案质量。
原文摘要 · Abstract (English)
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。