揭示音频与文本嵌入的语义鸿沟本质,提出无需训练的优化方法。
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

- 用PLS-SVD分解多模态嵌入,识别共享语义轴
- 仅保留少数可解释维度即显著缩小模态差距
- 零样本音频描述生成接近有监督效果,适合部署优化
对比语言-音频预训练(CLAP)模型广泛用于音频理解,并支持零样本条件替换。然而其性能受音频与文本嵌入间模态差距严重影响。现有解释多归因于锥形效应,视为均值偏移,但仅修正均值提升有限。信息不平衡与维度坍缩等假设亦被提出,但未在音频领域充分验证。尽管已有工作尝试将多模态对比嵌入分解为可解释概念,却未从概念分解视角分析模态差距。本文提出COMET(Concept space Organization and Modality gap Explanation with PLS-SVD Transformation),一种新颖的部分最小二乘奇异值分解(PLS-SVD)框架,揭示模态差距更广视角。该框架表明:仅少数可解释轴(捕捉共享概念)对相似度计算贡献显著,而均值分量仅部分反映模态差距。基于此,我们提出一种简单谱截断方法,无需训练即可缓解模态差距。该方法使零样本音频描述生成与条件替换逼近全监督性能,无需大型辅助内存或高成本计算。同时实现嵌入维度大幅压缩,且保持检索与音频描述任务强性能。
原文摘要 · Abstract (English)
Contrastive Language-Audio Pretraining (CLAP) models are widely used for audio understanding and support modality-agnostic condition swapping in many zero-shot applications. However, their performance is heavily affected by the modality gap between audio and text embeddings. Existing explanations mainly attribute this gap to the cone effect, treating it as a shift between mean embeddings, yet correcting the mean alone yields only limited improvements. Alternative hypotheses, such as information imbalance and dimensionality collapse, have also been proposed, but they remain insufficiently verified and have not been thoroughly studied in the audio domain. Meanwhile, several works attempt to decompose multimodal contrastive embeddings into interpretable concepts, but none explicitly analyze the modality gap from the perspective of concept decomposition. In this work, we introduce COMET (Concept space Organization and Modality gap Explanation with PLS-SVD Transformation), a novel partial least squares singular value decomposition (PLS-SVD) framework for CLAP that unveils a broader perspective of the modality gap. Our framework reveals that only a small, interpretable subset of axes, which captures shared concepts, contributes substantially to similarity computation, and that the mean component represents only partially the modality gap. Building on this insight, we propose a simple spectral truncation method that mitigates the modality gap in a training-free manner. The method enables zero-shot audio captioning with condition swapping to approach fully supervised performance, without requiring large auxiliary memory banks or expensive computation. At the same time, it achieves substantial embedding dimensionality reduction while preserving strong performance on retrieval and audio captioning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。