用结构化描述信息训练乳腺超声报告生成模型,少标注也能出好报告。
BUSTR: Descriptor-Aware Vision-Language Learning for Breast Ultrasound Report Generation
- 利用病变描述和影像特征自动生成报告,解决缺乏医生手写报告的问题。
- 在BrEaST和BUS-BRA数据集上提升报告相似度与描述恢复准确率。
- 适合医疗AI研发者和放射科医生,尤其关注报告自动化场景。
乳腺超声(BUS)报告依赖于临床有意义的病灶描述,包括BI-RADS分类、病灶形态、边界、回声强度、后方特征、病理及组织学信息。然而,多数公开的BUS数据集仅提供结构化标注和病灶掩码,缺少配对的放射科医生撰写报告,限制了视觉-语言模型在BUS报告生成中的发展。本文提出BUSTR,一种描述感知的视觉-语言框架,利用现有结构化病灶信息实现有限报告监督下的报告生成。BUSTR首先从可用标注和病灶掩码提取的影像组学特征中构建描述驱动的报告;然后使用多头Swin Transformer编码器,在部分重叠标注集的数据集上进行多任务监督,学习跨数据集的描述感知视觉表征;投影后的视觉令牌引导冻结的LLaMA语言模型,训练采用双层目标:词级交叉熵与表征级余弦对齐。推理时,BUSTR仅需输入超声图像,无需结构化描述、病灶掩码或影像组学特征即可生成报告。我们在公开的BrEaST和BUS-BRA数据集上评估,采用自然语言生成与临床有效性指标。结果表明,相比代表性基线模型,BUSTR在报告相似度与描述恢复能力上均有提升,尤其在病灶形态、边界、后方特征和病理类型方面表现突出,并在BrEaST上提升了BI-RADS分类的敏感性和F1分数。结果说明,即使无配对报告,结构化描述、病灶掩码与影像组学特征仍可为描述感知的乳腺超声报告生成提供有效监督。
原文摘要 · Abstract (English)
Breast ultrasound (BUS) reporting relies on clinically meaningful lesion descriptors, including BI-RADS category, lesion shape, margin, echogenicity, posterior features, pathology, and histology. However, many public BUS datasets provide structured annotations and lesion masks without paired radiologist-written reports, limiting the development of vision--language models for BUS report generation. We propose BUSTR, a descriptor-aware vision--language framework that uses structured lesion information to enable report generation under limited report supervision. BUSTR first constructs descriptor-derived reports from available annotations and radiomics features extracted from lesion masks. It then trains a multi-head Swin Transformer encoder with multitask supervision to learn descriptor-aware visual representations across datasets with partially overlapping annotation sets. The projected visual tokens condition a frozen LLaMA-based language model, and training is guided by a dual-level objective combining token-level cross-entropy with representation-level cosine alignment. At inference, BUSTR generates reports from BUS images without access to structured descriptors, lesion masks, or radiomics features. We evaluate BUSTR on the public BrEaST and BUS-BRA datasets using natural language generation and clinical efficacy metrics. BUSTR improves report similarity and descriptor recovery compared with representative report-generation baselines, with notable gains for lesion shape, margin, posterior features, and pathology, as well as improved BI-RADS sensitivity and F1-score on BrEaST. These results suggest that structured BUS descriptors, lesion masks, and radiomics features can provide useful supervision for descriptor-aware BUS report generation when paired radiologist-written reports are unavailable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。