arXiv:2512.00496cs.CLcs.AI2025-12

无需重训即可低成本接入新模态与100+语言,实现跨模态对齐。

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

  • 通过文本中心对齐机制,仅微调新模态数据即可融合多模态。
  • 在音频到文本检索中提升14.24个百分点,支持超100种语言。
  • 适合追求高效多模态多语言部署的开发者和研究者。

随着深度学习发展,单模态任务(如文本、图像、音频)正逐步转向多模态交互。现有模型通常需在多个模态上进行资源密集型训练,且扩展至新语言时也依赖同样高成本的流程。本文提出CACARA架构,采用涌现对齐学习,可在不重新训练整个模型的前提下,无缝集成新模态。该方法仅需用英语对齐数据微调新模态,便能从单语预训练中涌现出对100多种语言的支持,无需显式多语言预训练或文本编码器调优。该策略以接近单语模型的训练成本,保留原有知识,实现高达14.24个百分点的R@1音频到文本检索性能提升,显著优于现有先进模型,同时避免了跨模态与多语言的高额计算开销。

原文摘要 · Abstract (English)

As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimodal data. These connections bridge information gaps: an image can visually materialize a text, while audio can add context to an image. Researchers have developed numerous multimodal models, but most rely on resource-intensive training across multiple modalities. Similarly, extending these models to new languages often follows the same resource-heavy training strategy. In this work, we propose a multimodal and multilingual architecture, CACARA, trained through emergent alignment learning, enabling the seamless integration of new modalities into an existing bimodal/multimodal model without requiring full retraining. This work breaks new ground by demonstrating that this emergent alignment paradigm can unlock multilingual capabilities from monolingual training. By fine-tuning the newly incorporated modality only on data aligned with the English language, our model develops support for over 100 languages without explicit multilingual pretraining or tuning of the text encoder. Such emergent multimodal and multilingual properties are gained efficiently, preserving previously learned knowledge at a training cost comparable to that of a monolingual model. Our strategy achieves up to a 14.24 percentage points improvement in R@1 audio-to-text retrieval, outperforming state-of-the-art multimodal models -- all without the heavy computational cost of retraining across every modality and language.

多模态跨模态对齐多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。