arXiv:2508.04101cs.CV2025-08

提出双向交互框架,用正交约束提升医学视觉语言模型对齐效果

NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding

  • 设计双向交叉注意力适配器,动态生成跨模态查询增强交互
  • 引入正交正则化,分离新知识中的增量与全新成分,减少干扰
  • 仅新增146万参数,实测在肺炎数据集上提升2.3%性能

计算机辅助医学影像分析对疾病诊断和治疗规划至关重要。尽管视觉语言模型(如CLIP)具备强泛化能力,但其直接应用于医学影像仍受显著领域差距制约。现有方法如提示学习和单向模态交互通常仅将领域知识引入单一模态,未能充分利用CLIP的双模态结构,也忽略了双向跨模态交互的协同效应,导致模态对齐不足。本文提出NEARL(iNteracted quEry Adaptation with oRthogonaL Regularization),一种新型参数高效视觉语言模型框架,支持双向跨模态交互。NEARL包含两个核心组件:(1) 统一协同嵌入变换器(USEformer),动态生成紧凑的跨模态查询以促进交互;(2) 正交交叉注意力适配器(OCA),通过正交正则化将新知识解耦为真正新颖与增量成分。该设计降低增量成分干扰,使模型更专注学习新信息,改善模态交互。值得注意的是,NEARL仅引入146万可学习参数。在三种医学影像模态上的大量实验表明,其性能达到当前最优(如肺炎数据集相对提升2.3%),同时具备快速推理与低内存开销,凸显其在真实医疗视觉语言理解中的有效性。

原文摘要 · Abstract (English)

Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization ability, their direct application to medical imaging remains hindered by a substantial domain gap. Existing methods for bridging this gap, including prompt learning and unidirectional modality interaction, typically introduce domain knowledge into only one modality. However, such approaches fail to fully exploit CLIP's inherent dual-modality structure and overlook the synergistic effect of bidirectional cross-modal interaction, resulting in persistent modality misalignment. In this paper, we propose NEARL (iNteracted quEry Adaptation with oRthogonaL Regularization), a novel parameter-efficient VLM framework for bidirectional cross-modal interaction. NEARL consists of two key components: (1) the Unified Synergy Embedding Transformer (USEformer), which dynamically generates compact cross-modal queries to facilitate interaction; and (2) the Orthogonal Cross-Attention Adapter (OCA), which decouples new knowledge into truly novel and incremental components through orthogonal regularization. This design reduces interference from incremental components, enabling more focused learning of novel information and improving modality interaction in VLMs. Notably, NEARL introduces only 1.46M learnable parameters. Extensive experiments on three medical imaging modalities demonstrate state-of-the-art performance (e.g., a 2.3% relative improvement on the pneumonia dataset), along with fast inference and low memory overhead, highlighting its effectiveness for real-world medical vision-language understanding.

视觉语言医学影像跨模态正交约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。