arXiv:2505.08751cs.CLcs.CV2025-05被引 33

Aya Vision提升多语言多模态模型性能,解决跨语言对齐与遗忘问题。

Aya Vision: Advancing the Frontier of Multilingual Multimodality

  • 构建合成标注框架,生成高质量多语言多模态指令数据。
  • 32B模型超越90B级大模型,实现高精度跨模态生成。
  • 适合需要多语言能力的多模态应用开发者使用。

构建多模态语言模型面临核心挑战:视觉与语言模态对齐、高质量指令数据构建,以及引入视觉后文本能力退化。这些挑战在多语言场景中尤为突出,因多语言多模态数据稀缺、机器翻译易扭曲语义、灾难性遗忘更严重。为此,我们提出新型数据与建模技术:首先,开发合成标注框架,生成高质量、多样化的多语言多模态指令数据,使Aya Vision模型在多种语言下对多模态输入生成自然且受人类偏好的响应;其次,提出跨模态模型融合技术,有效缓解灾难性遗忘,同时保留文本生成能力并提升多模态生成性能。Aya-Vision-8B在多个强基线模型(如Qwen-2.5-VL-7B、Pixtral-12B、Llama-3.2-90B-Vision)中表现最佳;进一步扩展至Aya-Vision-32B,超越超过其两倍规模的模型(如Molmo-72B、LLaMA-3.2-90B-Vision)。本工作推动多语言多模态前沿进展,为低算力高效率高性能提供新思路。

原文摘要 · Abstract (English)

Building multimodal language models is fundamentally challenging: it requires aligning vision and language modalities, curating high-quality instruction data, and avoiding the degradation of existing text-only capabilities once vision is introduced. These difficulties are further magnified in the multilingual setting, where the need for multimodal data in different languages exacerbates existing data scarcity, machine translation often distorts meaning, and catastrophic forgetting is more pronounced. To address the aforementioned challenges, we introduce novel techniques spanning both data and modeling. First, we develop a synthetic annotation framework that curates high-quality, diverse multilingual multimodal instruction data, enabling Aya Vision models to produce natural, human-preferred responses to multimodal inputs across many languages. Complementing this, we propose a cross-modal model merging technique that mitigates catastrophic forgetting, effectively preserving text-only capabilities while simultaneously enhancing multimodal generative performance. Aya-Vision-8B achieves best-in-class performance compared to strong multimodal models such as Qwen-2.5-VL-7B, Pixtral-12B, and even much larger Llama-3.2-90B-Vision. We further scale this approach with Aya-Vision-32B, which outperforms models more than twice its size, such as Molmo-72B and LLaMA-3.2-90B-Vision. Our work advances multilingual progress on the multi-modal frontier, and provides insights into techniques that effectively bend the need for compute while delivering extremely high performance.

多模态多语言模型融合指令数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。