提出新方法诊断医疗视觉语言模型的模态失衡问题。
Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

- 用主成分基投影并加权计算方向性对齐分数,发现医学图像信息更丰富。
- SAS与医疗领域检索性能零标签相关性最强,优于现有所有指标。
- 适合关注医疗AI模型可解释性与临床部署的研究者使用。
视觉-语言模型(VLMs)在医疗图像-文本数据上表现不佳,但现有的诊断工具仍有限。现有表示对齐度量均为对称,将双模态信息合并为单一分数,掩盖了导致跨模态退化的主导模态。本文提出谱对齐得分(Spectral Alignment Score, SAS),一种非对称度量:将双模态投影至锚定模态的主特征基,并基于特征值加权计算各特征模式的相关性,生成方向性得分,其差值反映模态间信息不平衡。我们将SAS嵌入基准框架,评估15种VLMs在自然与医疗图像-文本数据集上的表现,对比6种对齐度量和双向检索任务。实验表明,医疗图像保留的结构信息显著多于对应临床报告,这种方向性不对称被所有现有度量忽略;且SAS在医疗领域与检索性能的零标签相关性最强,证明其作为临床部署实用诊断工具的潜力。代码已公开于https://github.com/iamalegambetti/medical-vlms-assessment。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existing representation alignment metrics are symmetric, collapsing both modalities into a single score and hiding which modality drives cross-modal degradation. We introduce the Spectral Alignment Score (SAS), an asymmetric metric that projects both modalities onto the principal eigenbasis of an anchor modality and computes eigenvalue-weighted per-eigenmode correlations, resulting in directional scores whose difference quantifies modality information imbalance. We embed SAS within a benchmarking framework evaluating 15 VLMs across natural and medical image-text datasets alongside 6 alignment metrics and bidirectional retrieval. Our experiments show that medical images retain richer structural information than their paired clinical reports, a directional asymmetry invisible to all competing metrics, and that SAS achieves the strongest zero-label correlation with retrieval performance in the medical domain, positioning it as a practical diagnostic tool for clinical deployment. Code is available at this URL: https://github.com/iamalegambetti/medical-vlms-assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。