无需人工标注,让医学影像模型自动理解图像位置与文本描述。
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

- 用大模型自动构建120万对医学影像与文本数据,实现无标注空间定位。
- 模型在报告生成、问答和定位任务上表现优秀,且加了定位训练不降低语言能力。
- 适合医学AI研究者,尤其关注影像与文本联合建模的场景。
我们研究如何在无需人工空间标注的情况下训练面向放射科的视觉-语言模型(VLM)。提出RefRad2D,一个包含120万张CT和MR图像-文本对的大规模双语(德/英)数据集,源自临床实践,通过大模型驱动的筛选与自动分割技术生成任务特定的视觉问答(VQA)和空间定位子集。基于该数据训练的RadGrounder模型可联合完成报告生成、视觉问答及边界框检测或分割形式的空间定位。在外部VQA基准(Slake、VQA-RAD)上,其表现优于专用医疗VLM;将我们的临床数据加入训练混合,显著提升开放性问题回答性能,证明数据集具备良好迁移能力。关键的是,增加空间监督并未降低语言质量,实现空间可验证输出而无需牺牲问答性能。
原文摘要 · Abstract (English)
We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs derived from clinical practice, with task-specific VQA and spatial grounding subsets generated automatically via LLM-based curation and automated segmentation. Trained on this data, our model RadGrounder jointly performs report generation, visual question answering, and spatial grounding via bounding-box detection or segmentation. On external VQA benchmarks (Slake, VQA-RAD), RadGrounder achieves competitive results with specialized medical VLMs. Adding our clinical data to the training mixture improves open-ended VQA over fine-tuning on the downstream datasets alone, showing the transferability of our dataset. Crucially, adding grounding supervision does not degrade language quality, enabling spatially verifiable outputs at no cost to VQA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。