让医学影像报告更准地定位病灶,支持多区域和非诊断性描述。
Generalised Medical Phrase Grounding

- 提出通用医学短语定位新任务,可返回零个、一个或多个区域。
- 在PadChest-GR和MS-CXR上对多区域和不可定位短语效果显著提升。
- 只需少量标注数据,且能与现有报告生成器无缝组合使用。
医学短语定位(MPG)将放射学报告中的文字描述映射到图像相应区域,使报告更易理解,尤其对非专业人士。现有系统多遵循指代表达理解(REC)范式,每条短语仅返回一个边界框,但真实报告常包含多区域发现、非诊断性文本及不可定位短语(如否定句或正常解剖描述)。为此,我们重新定义任务为通用医学短语定位(GMPG),每句话可对应零个、一个或多个带分值的图像区域。为此提出首个GMPG模型MedGrounder,采用两阶段训练:先在报告句子-解剖框对齐数据集上预训练,再在报告句子-人工标注框数据集上微调。在PadChest-GR和MS-CXR上的实验表明,MedGrounder具备强零样本迁移能力,在多区域和不可定位短语上优于传统REC方法及基于生成的基线模型,且所需人工标注框极少。最后证明,MedGrounder可与现有报告生成器组合,无需重训练即可生成带定位的报告。
原文摘要 · Abstract (English)
Medical phrase grounding (MPG) maps textual descriptions of radiological findings to corresponding image regions. These grounded reports are easier to interpret, especially for non-experts. Existing MPG systems mostly follow the referring expression comprehension (REC) paradigm and return exactly one bounding box per phrase. Real reports often violate this assumption. They contain multi-region findings, non-diagnostic text, and non-groundable phrases, such as negations or descriptions of normal anatomy. Motivated by this, we reformulate the task as generalised medical phrase grounding (GMPG), where each sentence is mapped to zero, one, or multiple scored regions. To realise this formulation, we introduce the first GMPG model: MedGrounder. We adopted a two-stage training regime: pre-training on report sentence--anatomy box alignment datasets and fine-tuning on report sentence--human annotated box datasets. Experiments on PadChest-GR and MS-CXR show that MedGrounder achieves strong zero-shot transfer and outperforms REC-style and grounded report generation baselines on multi-region and non-groundable phrases, while using far fewer human box annotations. Finally, we show that MedGrounder can be composed with existing report generators to produce grounded reports without retraining the generator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。