让AI生成报告像医生一样思考,提升乳腺钼靶报告准确性
Cross-Modal Clinical Knowledge Integration for Mammography Report Generation

- 按BI-RADS指南分步生成,融合患者历史影像与临床知识
- 在4个数据集上诊断类BI-RADS F1得分领先2.73%~3.27%
- 专设解析工具,可从自由文本提取结构化临床信息
乳腺癌是全球重大健康问题,钼靶筛查在早期发现中至关重要。大量筛查工作给放射科医生带来沉重负担,准确一致的报告生成成为关键临床挑战。现有自动化报告生成方法多聚焦于图像到文本的直接映射,忽视了放射科医生实际诊疗中的结构化推理流程。为此,我们提出MammoRG框架,通过遵循BI-RADS指南并融入既往临床知识,显式模拟临床报告流程以生成诊断报告。该框架采用两阶段训练:第一阶段基于分类监督,学习从四视图钼靶影像中整合相关临床先验知识;第二阶段引入术语感知的微调策略,将钼靶特异性临床术语作为原子语义单元建模,提升报告质量与临床一致性。为评估生成报告的临床效用,我们进一步开发了MammoRGTool,一款专门用于从自由文本报告中提取结构化临床信息的解析工具。大量实验表明,MammoRG在多个临床效能指标上持续优于现有方法,尤其在诊断相关BI-RADS F1指标上,分别在内部、外部1、外部2和VinDr-Mammo数据集上超越次优模型2.73%、2.04%、1.90%和3.27%。
原文摘要 · Abstract (English)
Breast cancer is a major global health concern, and mammography screening plays a central role in early detection. The large volume of screening examinations creates a substantial workload for radiologists, making accurate and consistent report generation a critical clinical challenge. Existing automated mammography report generation methods primarily focus on direct visual-to-text mapping, while overlooking the structured clinical reasoning process followed by radiologists in real-world practice. To address this limitation, we propose MammoRG, a mammography report generation framework that explicitly simulates the clinical reporting workflow by following the BI-RADS guideline and incorporating prior clinical knowledge to produce diagnostic reports. Specifically, MammoRG adopts a two-stage training framework. In the first stage, the model learns to integrate clinically relevant prior knowledge from a patient's four-view mammograms through classification-based supervision. In the second stage, a terminology-aware supervised fine-tuning strategy is introduced to model mammography-specific clinical terms as atomic semantic units, enabling the generation of high-quality reports with improved clinical consistency. To facilitate clinical efficacy evaluation of generated reports, we further develop MammoRGTool, a dedicated mammography report parsing tool that extracts structured clinical information from free-text reports. Extensive experiments demonstrate that MammoRG consistently outperforms existing methods across multiple clinical efficacy metrics, particularly in diagnosis-related BI-RADS F1, where it surpasses the second-best model by 2.73%, 2.04%, 1.90%, and 3.27% on the internal, external 1, external 2, and VinDr-Mammo datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。