arXiv:2603.01096cs.CVcs.AI2026-03

将视觉与语言嵌入对齐到统一空间,实现多语言零样本理解与生成。

Unified Vision-Language Modeling via Concept Space Alignment

  • 通过后处理对齐将视觉编码映射到文本嵌入空间,构建统一表征
  • 在视频描述任务中超越现有模型,低资源语言性能提升显著
  • 支持零样本跨模态概念理解,适合多语言多模态应用开发

我们提出V-SONAR,基于纯文本嵌入空间SONAR扩展的视觉-语言嵌入空间,支持1500种文本语言和177种语音语言。通过后处理对齐流程,将现有视觉编码器的表示映射至SONAR空间。实验表明,V-SONAR在文本到视频检索任务中表现优异;结合OMNISONAR文本解码器,在视频字幕任务上超越当前最优模型:DREAM-1K(BLEU 23.9 vs. 19.6)、PE-VIDEO(BLEU 39.0 vs. 30.0)。利用V-SONAR,我们首次证明仅用英语训练的大型概念模型(LCM)可零样本理解单/多视觉概念。进一步提出V-LCM,通过视觉-语言指令微调扩展LCM,采用与原版相同的潜在扩散目标进行下一步嵌入预测训练。大规模多语言、多模态指令数据实验证明:V-LCM在图像/视频字幕与问答任务上达到顶尖水平,且在61/62种语言(从丰富到低资源)中显著优于现有模型。

原文摘要 · Abstract (English)

We introduce V-SONAR, a vision-language embedding space extended from the text-only embedding space SONAR (Omnilingual Embeddings Team et al., 2026), which supports 1500 text languages and 177 speech languages. To construct V-SONAR, we propose a post-hoc alignment pipeline that maps the representations of an existing vision encoder into the SONAR space. We thoroughly evaluate V-SONAR and show that its embeddings achieve competitive performance on text-to-video retrieval. Equipped with the OMNISONAR text decoder, V-SONAR further surpasses state-of-the-art vision-language models on video captioning tasks, including DREAM-1K (BLEU 23.9 vs. 19.6) and PE-VIDEO (BLEU 39.0 vs. 30.0). Leveraging V-SONAR, we first demonstrate that the Large Concept Model (LCM; LCM team et al. 2024) operating in SONAR and trained with English text only, can perform both single- and multi-visual concept understanding in a zero-shot manner. Finally, we introduce V-LCM, which extends the LCM with vision-language instruction tuning. V-LCM encodes vision and language inputs into an unified sequence of latent embeddings via V-SONAR and SONAR, and it is trained with the same latent diffusion objective for next-embedding prediction as in LCM's text-only pre-training. Experiments on a large-scale multilingual and -modal instruction-tuning data mixture highlight the potential of V-LCM: V-LCM matches state-of-the-art vision-language models on tasks covering image/video captioning and question answering, while significantly outperforming them across 61 rich- to low-resource languages out of all 62 tested languages.

多模态统一表征零样本多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。