arXiv:2502.15392cs.AIcs.CL2025-02被引 13

面向10种印度语言的多模态模型,让低资源语言也能用上先进AI。

Chitrarth: Bridging Vision and Language for a Billion People

  • 用多语言大模型+视觉模块,融合10种印度语言图文数据训练
  • 在低资源语言任务上达到当前最优表现,英语性能也不下降
  • 开源评估框架BharatBench,推动多语言多模态研究

当前多模态基础模型主要基于英语或高资源欧洲语言数据训练,难以适配其他中低资源语言。为此,我们提出Chitrarth(Chitra: 图像;Artha: 含义),一种面向10种主流印度语言的包容性视觉-语言模型,旨在提升跨语言视觉推理能力。模型通过整合最先进的多语言大语言模型与视觉模块,在多语言图文数据上进行训练。同时,我们构建了BharatBench——一个全面评估多模态模型在多种印度语言上表现的基准框架。实验表明,该模型在低资源语言任务中取得当前最优结果,且保持英语性能优势。本研究旨在建立多语言多模态新标准,为未来多样化AI系统发展奠定基础。

原文摘要 · Abstract (English)

Recent multimodal foundation models are primarily trained on English or high resource European language data, which hinders their applicability to other medium and low-resource languages. To address this limitation, we introduce Chitrarth (Chitra: Image; Artha: Meaning), an inclusive Vision-Language Model (VLM), specifically targeting the rich linguistic diversity and visual reasoning across 10 prominent Indian languages. Our model effectively integrates a state-of-the-art (SOTA) multilingual Large Language Model (LLM) with a vision module, primarily trained on multilingual image-text data. Furthermore, we also introduce BharatBench, a comprehensive framework for evaluating VLMs across various Indian languages, ultimately contributing to more diverse and effective AI systems. Our model achieves SOTA results for benchmarks across low resource languages while retaining its efficiency in English. Through our research, we aim to set new benchmarks in multilingual-multimodal capabilities, offering substantial improvements over existing models and establishing a foundation to facilitate future advancements in this arena.

多模态多语言视觉语言印度语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。