构建结构化多模态数据集,提升胸部X光报告生成的临床准确性。
MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation
- 将解剖区域、文本描述、图像坐标和报告目标整合为统一证据单元
- 使用多模态证据训练模型,在临床指标上优于纯文本或边界框方法
- 适配后性能媲美开源大模型,适合医学影像与AI交叉研究者
胸部X光报告生成遵循基于解剖区域的临床工作流程,放射科医生需逐区域检查并整合局部发现。然而,现有资源中的监督信号以碎片化形式存在。我们提出MMRad-22K,一个将区域文本观察、解剖定位坐标、局部图像证据与报告目标组织成结构化多模态证据单元的数据集。为验证该设计,我们对比了不同证据形式的生成效果,发现结构化多模态证据普遍优于仅文本或基于边界框的证据。随后,我们使用MMRad-22K适配统一的LVLM主干模型,结果显示,基于多模态证据的微调在语言和临床相关指标上均优于仅文本证据微调与端到端训练。在相同评估协议下,该模型表现可比肩多个开源LVLM基准。结果表明,MMRad-22K是符合临床读片流程的实用多模态资源,适用于训练与评估胸部X光报告生成任务。
原文摘要 · Abstract (English)
Chest X-ray (CXR) reporting follows a region-based clinical workflow in which radiologists inspect anatomical regions and integrate localized findings into a final report. However, existing resources for CXR report generation provide these supervision signals in fragmented forms. We introduce MMRad-22K, a dataset that organizes regional textual observations, anatomical grounding coordinates, localized image evidence, and report targets into structured multimodal evidence units for CXR report generation. To motivate this formulation, we first compare different evidence formats for report generation and find that structured multimodal evidence is generally more useful than text-only or bounding box-based evidence. We then adapt a unified LVLM backbone using MMRad-22K and show that adaptation with multimodal evidence outperforms both textual-evidence adaptation and end-to-end adaptation on language and clinically oriented metrics. Under the same evaluation protocol, the adapted model also reaches a performance level comparable to several open-source LVLM references. Together, these results support MMRad-22K as a practical structured multimodal resource for training and evaluating CXR report generation aligned with clinical reading workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。