用自由文本提炼医学影像细粒度知识,提升结构化报告生成准确率
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
- 通过大模型从海量自由文本中提取细粒度医学信息,构建视觉原型知识库
- 在Rad-ReStruct上实现当前最佳性能,对细节属性预测提升显著
- 适合需要高精度结构化报告的医疗AI研发人员参考
结构化放射科报告可实现更快速、一致的沟通,但自动化仍具挑战,因模型需在有限结构化监督下对罕见发现和属性做出大量细粒度决策。相比之下,自由文本报告在日常诊疗中规模巨大,通过详细描述隐式包含与图像相关的细粒度信息。为利用这种非结构化知识,我们提出ProtoSR,一种将自由文本信息注入结构化报告生成的方法。首先,设计自动提取流程,使用指令微调的大语言模型挖掘超过8万例MIMIC-CXR研究,构建与结构化报告模板对齐的多模态知识库,每个答案选项由一个视觉原型表示。基于该知识库,ProtoSR训练时检索与当前图像-问题对相关的原型,并通过原型条件残差增强模型预测,提供数据驱动的第二意见,选择性纠正错误。在Rad-ReStruct基准测试中,ProtoSR达到当前最优结果,尤其在细节属性问题上改善最明显,证明了整合自由文本信号对细粒度图像理解的价值。
原文摘要 · Abstract (English)
Structured radiology reporting promises faster, more consistent communication than free text, but automation remains difficult as models must make many fine-grained, discrete decisions about rare findings and attributes from limited structured supervision. In contrast, free-text reports are produced at scale in routine care and implicitly encode fine-grained, image-linked information through detailed descriptions. To leverage this unstructured knowledge, we propose ProtoSR, an approach for injecting free-text information into structured report population. First, we introduce an automatic extraction pipeline that uses an instruction-tuned LLM to mine 80k+ MIMIC-CXR studies and build a multimodal knowledge base aligned with a structured reporting template, representing each answer option with a visual prototype. Using this knowledge base, ProtoSR is trained to retrieve prototypes relevant for the current image-question pair and augment the model predictions through a prototype-conditioned residual, providing a data-driven second opinion that selectively corrects predictions. On the Rad-ReStruct benchmark, ProtoSR achieves state-of-the-art results, with the largest improvements on detailed attribute questions, demonstrating the value of integrating free-text derived signal for fine-grained image understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。