arXiv:2506.10416cs.MMcs.SD2025-06被引 2

用声音替代视觉,探索跨模态对齐极限下的模型表现

Can Sound Replace Vision in LLaVA With Token Substitution?

  • 构建细粒度音画对齐数据集,实现超精准对齐训练
  • 图像主导模型检索强但生成差,文本主导模型平衡更好
  • 揭示模型架构决定对齐策略的适用性,适合跨模态研究者

当音画对齐被推向极致时会发生什么?现有数据集将对齐视为二元状态(同步或不同步),无法揭示其内在连续性。为此,我们构建了一个包含细粒度对齐评分的综合性数据集,揭示了音画感知对应关系的隐藏光谱。基于这些精确评分,我们仅用完全匹配的音画对齐样本训练出“超对齐”表征,并系统研究这种极端对齐如何影响跨模态检索与生成任务中的模型行为。研究涵盖两类编码器:以图像为中心的编码器(预训练时以视觉为媒介连接模态)和以文本为中心的编码器(直接进行音-语对齐预训练)。首先评估其在跨模态检索和视觉-语言描述生成任务上的基线性能,随后使用高度一致的音画数据将所有编码器重对齐至CLIP空间,观察性能变化。结果表明,初始架构类型决定了编码器对对齐过程的响应:图像主导编码器在跨模态检索中表现优异,但高强度对齐导致语言信息压缩,降低描述生成质量;而文本主导编码器具备更强的语言真实性,能在两个目标间保持更好平衡。

原文摘要 · Abstract (English)

What happens when we push audio-visual alignment to its absolute limits? To systematically investigate this question, we needed datasets with granular alignment quality annotations, but existing datasets treat alignment as binary, either synchronized or not. To address this limitation, we developed a comprehensive dataset featuring detailed alignment scores that reveal the hidden spectrum of audio-visual perceptual correspondence. Using these precise scores, we create "superaligned" representations by training exclusively on the most perfectly matched audio-visual pairs, then conduct our systematic investigation into how this extreme alignment transforms perceptual model behavior across retrieval and generation tasks. The encoders under study fall into two main groups consisting of image-centric encoders that were pretrained using visual modalities as intermediary hubs for connecting modalities, and text-centric encoders that were pretrained with direct audio-language alignment. We first measure the baseline performance of these encoders on two key tasks, namely cross-modal retrieval and text description generation in vision-language models. Subsequently, we realign all encoders with the CLIP space using highly coherent audio-visual data and observe the performance changes. Our findings reveal that the initial architectural type of the encoder determines how it responds to the alignment process. Image-centric encoders, which are inherently designed for alignment, demonstrate exceptional performance in cross-modal retrieval, but this intensive alignment causes compression of unique linguistic information and reduces the quality of their text description generation in vision-language models. In contrast, text-centric encoders, which possess stronger linguistic authenticity, are able to maintain a better balance between the two objectives.

跨模态音画对齐视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。