用密集临床对比监督,让医学大模型更准地更新知识
Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models

- 用四种不同格式的监督数据对比知识更新效果
- 采用EMQ格式时模型在癌症分期任务上准确率达64.8%
- 密集临床对比能增强模型对新旧知识的区分能力
医学知识持续更新,使大语言模型易依赖过时但看似合理的信息。我们研究在相同训练预算下,监督格式如何影响医学知识更新。提出SEER-Bench,一个基于最新版SEER研究数据的时序锚定肿瘤分期基准,并将NCCN肿瘤指南中的更新事件转化为四种监督格式:EMQ、MSQ、FITB和SAQ。在SEER-Bench和HealthBench Professional上,EMQ在相同预算的SFT方法中表现最稳定,具有最佳外部迁移与保留能力。使用EMQ监督的4B模型在时序锚定肿瘤分期任务中达到64.8%答案准确率和59.6%推理准确率。诊断分析表明,EMQ暴露更密集的临床对比信号,同时保持判别性表示且对基础模型的偏移更小。结果表明,医学知识更新不仅依赖更新算法,还取决于知识的监督结构。
原文摘要 · Abstract (English)
Medical knowledge changes continually, making large language models vulnerable to relying on outdated yet clinically plausible information. We study whether the format of supervision affects medical knowledge updating under a matched training-budget setting. We introduce SEER-Bench, a temporally anchored oncology-staging benchmark curated from the latest versioned SEER Research Data release, and render identical medical update events from NCCN oncology guidelines into four supervision formats: EMQ, MSQ, FITB, and SAQ. Across SEER-Bench and HealthBench Professional, EMQ gives the most stable external transfer and retention among same-budget SFT variants. With EMQ supervision, the updated 4B model produces competitive results on temporally anchored oncology staging, reaching 64.8% answer accuracy and 59.6% rationale accuracy on SEER-Bench. Diagnostic analyses suggest that EMQ exposes denser clinical contrast signals while preserving discriminative representations with smaller movement from the base model. These results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。