arXiv:2510.05410cs.CLcs.LG2025-10被引 1

用偏好优化让小模型写出更专业的重症心衰护理记录。

Aligning Language Models with Clinical Expertise: DPO for Heart Failure Nursing Documentation in Critical Care

  • 用专家标注的对比数据直接优化模型,提升文本质量。
  • 模型生成文档的准确性和完整性提升超14分,流畅度显著改善。
  • 适合医疗AI落地,兼顾隐私与效率,可部署于医院系统。

重症监护室(ICU)护理记录包含关键临床信息,但常因术语不统一、风格随意而影响质量,尤其在心衰管理中尤为突出。本研究采用直接偏好优化(DPO)方法,基于MIMIC-III数据库中的8,838份心衰护理记录及21,210组由专家验证的GPT输出、模型生成与原始记录构成的偏好对,微调可本地部署的Mistral-7B语言模型。评估显示,经优化后,BLEU得分从0.173升至0.318(+84%),BERTScore从0.828增至0.891(+7.6%),专家评分在准确性、完整性、逻辑一致性、可读性与结构清晰度上分别提升14.4、14.5、14.1、11.1和6.0分。结果表明,DPO能有效将轻量级临床语言模型对齐专家标准,支持隐私保护下的电子病历AI辅助书写,降低行政负担并提升患者安全。

原文摘要 · Abstract (English)

Nursing documentation in intensive care units (ICUs) provides essential clinical intelligence but often suffers from inconsistent terminology, informal styles, and lack of standardization, challenges that are particularly critical in heart failure care. This study applies Direct Preference Optimization (DPO) to adapt Mistral-7B, a locally deployable language model, using 8,838 heart failure nursing notes from the MIMIC-III database and 21,210 preference pairs derived from expert-verified GPT outputs, model generations, and original notes. Evaluation across BLEU, ROUGE, BERTScore, Perplexity, and expert qualitative assessments demonstrates that DPO markedly enhances documentation quality. Specifically, BLEU increased by 84% (0.173 to 0.318), BERTScore improved by 7.6% (0.828 to 0.891), and expert ratings rose across accuracy (+14.4 points), completeness (+14.5 points), logical consistency (+14.1 points), readability (+11.1 points), and structural clarity (+6.0 points). These results indicate that DPO can align lightweight clinical language models with expert standards, supporting privacy-preserving, AI-assisted documentation within electronic health record systems to reduce administrative burden and improve ICU patient safety.

医疗AI偏好优化护理记录轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。