用改进的文本编码器提升胸部X光图像与报告的跨模态检索效果
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
- 针对放射报告风格差异,设计领域自适应语言模型编码器
- 在160万对数据上实现MIMIC-CXR绿分0.308、Open-I 0.618
- 特别适合处理含缩写和仅印象的医院报告,提升泛化能力
医学影像与临床文本的多模态学习是医疗数据驱动信息学的核心挑战,有效的跨模态对齐对可扩展分析与检索至关重要。在胸片领域,视觉-语言预训练受限于报告内容不一,包括缩写、仅含印象的条目及机构特有写作风格。与通用场景不同,直接合并大量噪声报告可能导致多模态学习性能停滞甚至下降。本文提出一种领域适配的双向大语言模型文本编码器,通过掩码词预测与监督对比学习,在风格多样但临床等价的报告变体上训练,生成稳健且可泛化的文本嵌入。随后将其集成到双塔对比框架中,采用参数高效适配方法增强图像-文本对齐。在来自公开数据集和去标识化医院队列的160万对数据上,所提模型显著提升双向检索准确率与外部泛化能力,分别在MIMIC-CXR上取得0.308的GREEN得分,在Open-I上达0.618,同时有效缓解了含缩写的、仅印象的医院报告加入训练后导致的性能退化。
原文摘要 · Abstract (English)
Multimodal learning from paired medical images and clinical text is a central challenge in medical data-driven informatics, where effective cross-modal alignment is critical for scalable analysis and retrieval. In chest radiography, vision-language pretraining is constrained by heterogeneous radiology reports that contain abbreviations, impression-only notes, and institution-specific writing styles. Unlike general-domain settings, naively aggregating large collections of noisy reports can plateau or even degrade multimodal learning when reporting styles differ substantially. We propose a domain-adapted bidirectional large language model text encoder for chest radiograph reports, trained with masked token prediction and supervised contrastive learning on stylistically diverse but clinically equivalent report variants to produce robust, generalizable text embeddings. We then integrate this encoder into a dual-tower contrastive vision-language framework using parameter-efficient adaptation to improve image-text alignment. Across 1.6 million paired studies from public datasets and a de-identified hospital cohort, the proposed models improve bidirectional retrieval accuracy and external generalization, achieving GREEN scores of 0.308 on MIMIC-CXR and 0.618 on Open-I, while reducing the degradation observed when abbreviation-rich, impression-only hospital reports are added to training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。