arXiv:2605.17152cs.CL2026-05

让多语言多模态模型在资源匮乏时仍能高效运行

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

  • 用轻量数据构建与适配器对齐三模态信息
  • 在低资源下实现跨文本、语音、视觉的多语言模型
  • 适合关注低资源多模态应用的研究者与工程师

多模态大模型正从视觉-语言向视听读三模态演进,但现有流程与评估基准仍以英语为中心且计算成本高。本教程面向低资源语言场景,综述多语言多模态研究基础,涵盖近期模型(PALO、Maya)及语音-文本大模型。内容包括低成本数据创建与整理方法、三模态对齐的适配器架构、超越英语的文化敏感评估,以及如何微调紧凑型多语言视觉语言模型、搭建语音→文本→大模型处理链路。教程为互动式半日课程,面向致力于多语言多模态AI在低资源环境下落地的研究人员与实践者。

原文摘要 · Abstract (English)

Multimodal LLMs are evolving from vision-language to tri-modality that see, hear, and read, yet pipelines and benchmarks remain English-centric and compute-heavy. The tutorial offers an overview of this emerging research area for multilingual multimodality across text, speech, and vision under limited data/compute budgets, synthesizing foundations, recent multilingual models (PALO, Maya), speech-text LLMs. We cover low-cost data creation/curation; adapter stacks for tri-modal alignment; culture-aware evaluation beyond English and hands on resources for fine-tuning a compact multilingual VLM and wiring a speech->text->LLM pipeline. The content will be delivered as an interactive half-day tutorial, designed for researchers and practitioners working on multilingual, multimodal AI in low-resource language settings.

多模态低资源多语言语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。