arXiv:2609.01757cs.CV2026-09

用医学概念空间和分块特征融合,提升胸部X光零样本分类精度。

AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $\alpha$-Corrected Binary Cross Entropy and Factorized Latent Supervision

论文配图:AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $\alpha$-Corrected Binary Cross Entropy and Factorized Latent Supervision
图 1 · 摘自论文原文
  • 基于大语言模型构建医学概念空间,减少对比学习噪声。
  • 在16个基准上达最优平均性能,7个定位任务和3个分割任务领先。
  • 无需增加参数的轻量级融合模块,适合医疗视觉多模态研究。

视觉-语言预训练模型为开放词汇的胸部放射学理解提供了可扩展路径,但两个方面仍待深入:如何利用从医学报告中提取的结构化临床语义减少对比学习中的批内噪声,以及如何设计跨模态融合以实现更准确的空间定位且不增加复杂度。本文提出AlphaRAD,通过两项贡献解决上述问题。首先,利用大语言模型解析医学报告构建大规模结构化医学概念空间用于训练,从而缓解批内学习噪声,消除对比学习中的启发式配对过程,使AlphaRAD自然成为通过α校正二元交叉熵训练的医学概念判别器。其次,提出FLaS(因子化潜在监督),一种极简而高效的跨模态特征融合模块,将视觉-语言预训练模型表示分解为独立子空间,并使用专用对齐监督增强空间定位表达能力,且不引入额外参数。大量实证验证表明,AlphaRAD在多种胸部放射学任务中展现出强大的零样本泛化能力。显著地,在16个分类基准上达到当前最佳平均性能,同时在7个定位/短语定位和3个分割数据集上取得各自最佳结果。

原文摘要 · Abstract (English)

Vision-Language Pretrained Models (VLPMs) offer a scalable path to open-vocabulary chest radiology understanding, yet two aspects remain underexplored: how structured clinical semantics extracted from medical reports can reduce in-batch noise during contrastive learning, and how cross-modal fusion can be designed to produce more faithful spatial grounding without added complexity. We introduce AlphaRAD, addressing these opportunities through two contributions. First, we construct a large-scale structured medical concept space from medical reports parsed by a Large Language Model for training, thereby mitigating in-batch learning noise and removing heuristic pair matching in contrastive learning, and thus naturally positioning AlphaRAD as a medical concept discriminator trained via $\alpha$-Corrected Binary Cross-Entropy. Second, we propose FLaS (Factorized Latent Supervision), an extremely simple yet effective cross-modal feature fusion module that factorizes VLPM representations into independent subspaces, using dedicated alignment supervision to enhance the expressiveness of spatial grounding without introducing additional model parameters. Through extensive empirical validation, AlphaRAD shows strong zero-shot generalization across diverse chest radiology tasks. Notably, it establishes state-of-the-art average performance across 16 classification benchmarks, while achieving individual state-of-the-art results via distinct gains on 7 grounding/phrase grounding and 3 segmentation datasets.

医学影像零样本跨模态视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。