arXiv:2601.15408cs.CVcs.AI2026-01中稿 · CVPR被引 1

CURE通过课程学习提升医学报告生成的视觉定位与事实一致性

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

  • 基于模型表现动态调整难样本采样,强化图文对齐
  • 视觉定位准确率提升0.35 IoU,报告质量增0.192 CXRFEScore
  • 无需额外数据,适合医疗AI报告系统开发

医学视觉语言模型能自动生成放射科报告,但常在视觉定位和事实一致性上出错,导致文本描述与影像证据不符。我们提出CURE,一种错误感知的课程学习框架,在不增加数据的前提下提升定位精度与报告可靠性。CURE在公开数据集上微调多模态指令模型,涵盖短语定位、基于视觉的报告生成及解剖结构定位报告生成任务。方法根据模型表现动态调整采样策略,重点训练更难样本,以改善空间与文本对齐。实验显示,该方法使定位准确率提升0.35 IoU,报告质量提高0.192 CXRFEScore,幻觉现象减少18.6%。CURE是一种数据高效框架,显著提升生成结果的准确性与可信度。代码已开源于https://github.com/PabloMessina/CURE,模型权重可在https://huggingface.co/pamessina/medgemma-4b-it-cure 获取。

原文摘要 · Abstract (English)

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present CURE, an error-aware curriculum learning framework that improves grounding and report quality without any additional data. CURE fine-tunes a multimodal instructional model on phrase grounding, grounded report generation, and anatomy-grounded report generation using public datasets. The method dynamically adjusts sampling based on model performance, emphasizing harder samples to improve spatial and textual alignment. CURE improves grounding accuracy by +0.35 IoU, boosts report quality by +0.192 CXRFEScore, and reduces hallucinations by 18.6%. CURE is a data-efficient framework that enhances both grounding accuracy and report reliability. Code is available at https://github.com/PabloMessina/CURE and model weights at https://huggingface.co/pamessina/medgemma-4b-it-cure

医学报告生成视觉定位课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。