arXiv:2605.24902cs.CLcs.AI2026-05

强推理模型生成临床病历反不如弱模型,需谨慎评估。

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

  • 控制变量测试不同模型在推理与源文档检索下的表现
  • 非推理版GPT-5.4在三数据集上综合质量最高
  • 强推理反而降低GPT-5.4表现,适合医疗文书生成的模型需专项评估

具备推理能力的大语言模型在医学推理基准上表现优异,但其性能是否适用于结构化临床文档生成仍不明确。本研究基于涵盖OMI Health、ACI-Bench和PriMock57的源感知基准,考察了GPT-5.4、DeepSeek-V4-Flash和Gemma-4-E4B在临床对话转写为SOAP病历任务中的表现。采用2×2受控实验设计,独立调节医生原生推理与同源检索增强生成(RAG)。通过七项自动指标及两名参考感知的LLM裁判进行评估。两种评估方式均显示,非推理版本的GPT-5.4整体质量最优,而深度求索-V4-Flash在开启推理的配置中表现最佳。启用推理显著降低了GPT-5.4在全部三个数据集上的表现,而同源RAG仅带来较小且模型依赖的提升。总体表明,不能默认更强推理能力能提升对精度敏感的SOAP病历生成,必须开展专门的任务评估。

原文摘要 · Abstract (English)

Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical documentation; we investigate this question using SOAP note generation from clinical dialogue in a source-aware benchmark spanning OMI Health, ACI-Bench, and PriMock57. We evaluate GPT-5.4, DeepSeek-V4-Flash, and Gemma-4-E4B in a controlled 2x2 design that independently toggles provider-native reasoning and same-source retrieval-augmented generation (RAG). Outputs are assessed using seven automatic metrics alongside two reference-aware LLM judges. Both evaluation approaches agree that a non-reasoning GPT-5.4 configuration achieves the highest overall quality, while DeepSeek-V4-Flash performs best among reasoning-enabled configurations. Enabling reasoning significantly degrades GPT-5.4 performance across all three datasets, whereas same-source RAG yields smaller, model-dependent improvements. Overall, the findings indicate that stronger reasoning capability should not be assumed to improve fidelity-sensitive SOAP note generation without dedicated, task-specific evaluation.

临床生成大模型评估SOAP病历推理陷阱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。