无需医生反馈,用自动评分提升肺部X光报告生成模型
CheXalign: Preference fine-tuning in chest X-ray interpretation models without human feedback
- 用公开报告数据和自动评分替代人工反馈进行偏好微调
- 在MIMIC-CXR上达到当前最优的CheXbert得分
- 适合医疗AI研发者和想降低标注成本的团队
放射科医生在将医学影像转化为可行动报告方面起关键作用,但该领域面临人员短缺和工作量增加的挑战。尽管基于视觉语言模型(VLM)的自动化方法展现出作为辅助工具的潜力,但其准确率要求极高。目前多数放射学VLM仅依赖监督微调,而通用领域中偏好微调已成为标准流程。然而,在放射学中大规模获取放射科医生反馈成本过高。为此,我们提出一种自动化偏好反馈管道,专注于胸部X光片报告生成(RRG)。该方法利用包含图像与放射科医生撰写的参考报告的公开数据集,结合基于参考的评估指标(即“法官”),无需额外医生反馈。我们研究了长度滥用导致的奖励过优化问题,并引入长度可控的GREEN分数。最佳设置在MIMIC-CXR数据集上实现当前最优的CheXbert得分,同时在六个额外图像感知与推理任务中保持平均稳健表现。
原文摘要 · Abstract (English)
Radiologists play a crucial role in translating medical images into actionable reports. However, the field faces staffing shortages and increasing workloads. While automated approaches using vision-language models (VLMs) show promise as assistants, they require exceptionally high accuracy. Most current VLMs in radiology rely solely on supervised fine-tuning. Meanwhile, additional preference fine-tuning in the post-training pipeline has become standard practice in the general domain. The challenge in radiology lies in the prohibitive cost of obtaining radiologist feedback at scale. To address this challenge, we propose an automated pipeline for preference feedback, focusing on chest X-ray radiology report generation (RRG). Specifically, our method leverages publicly available datasets containing pairs of images and radiologist-written reference reports with reference-based metrics, or Judges, eliminating the need for additional radiologist feedback. We investigate reward overoptimization via length exploitation in this setting and introduce a length-controlled version of the GREEN score. Our best-performing setup achieves state-of-the-art CheXbert scores on the MIMIC-CXR dataset for the RRG task while on average maintaining robust performance across six additional image perception and reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。