用强化学习提升医学影像报告分类准确率与推理能力
Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports
- 先微调模型再用GRPO优化预测格式与准确率
- 在三个数据集上分类准确率进一步提升,推理更完整
- 适合需要高可靠性和可解释性的医疗AI应用
从放射科报告中准确分类疾病对多种应用场景至关重要。尽管轻量级大模型的监督微调(SFT)能提升准确率,但可能损害推理能力。本文提出两阶段方法:先在疾病标签上进行SFT,再通过组相对策略优化(GRPO)在无推理监督条件下优化准确率与输出格式。在三个由放射科医生标注的数据集上,SFT表现优于基线,而GRPO进一步提升了分类性能,并增强了推理的召回率与完整性。
原文摘要 · Abstract (English)
Accurate disease classification from radiology reports is essential for many applications. While supervised fine-tuning (SFT) of lightweight LLMs improves accuracy, it can degrade reasoning. We propose a two-stage approach: SFT on disease labels followed by Group Relative Policy Optimization (GRPO) to refine predictions by optimizing accuracy and format without reasoning supervision. Across three radiologist-annotated datasets, SFT outperformed baselines and GRPO further improved classification and enhanced reasoning recall and comprehensiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。