让放射科报告生成更贴近临床实际,结合解剖位置与病程变化。
Spatio-Temporal and Clinical Conditioning for Fine-Grained Radiology Report Retrieval

- 通过检测解剖区域,融合当前临床信息和前后影像对比变化进行检索。
- 在MIMIC-CXR数据集上,检索与临床指标均优于现有方法。
- 适合需要精准、动态病程描述的医疗AI研发人员使用。
放射科对现代医疗至关重要,但影像检查量上升与人力资源短缺加剧了报告负担与临床工作压力。自动化放射科报告生成有望缓解这一问题,但现有基于检索的方法仍存在僵化、缺乏明确解剖定位,且未考虑纵向疾病进展与可用临床背景的问题。本文提出STAR3,一种多模态、时空感知的注意力检索框架,用于放射科报告生成,可将区域级解剖信息与临床指征及胸部X光片间的纵向变化对齐。该框架采用目标检测器识别有意义的解剖区域,并基于当前临床背景及前后检查间的变化,检索语义相关的报告句子。这种设计使报告生成更符合临床实践中的解剖与时间依赖性。在MIMIC-CXR数据集上的实验表明,STAR3在检索、NLP与临床指标上均优于现有方法,验证了结合解剖、时间与临床条件进行检索对推进自动化放射科报告生成的价值。
原文摘要 · Abstract (English)
Radiology is vital to modern healthcare, but rising imaging demand and persistent workforce shortages strain reporting capacity and clinical workflows. Automated radiology report generation has the potential to support radiologists and help alleviate this burden; however, existing retrieval-based methods remain rigid, lack explicit anatomical grounding, and do not account for longitudinal disease progression or available clinical context. In this work, we introduce STAR3, a multimodal, spatio-temporal, attentive retrieval framework for radiology report generation that aligns region-level anatomical information with clinical indications and longitudinal changes across chest X-ray studies. Our framework employs an object detector to identify anatomically meaningful regions and retrieves semantically relevant report sentences conditioned on both current clinical context and changes observed between prior and current examinations. This design enables anatomically and temporally grounded report generation that better reflects clinical reporting practice. Experiments on the MIMIC-CXR dataset demonstrate that STAR3 outperforms current retrieval-based approaches on retrieval, NLP and clinical metrics, highlighting the value of conditioning retrieval anatomically, temporally and clinically for advancing automated radiology report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。