arXiv:2410.17532cs.CLcs.AI2024-10综述被引 5

系统梳理多语言大模型开发与应用,助力AI跨越语言鸿沟。

Responsible Multilingual Large Language Models: A Survey of Development, Applications, and Societal Impact

  • 构建从数据到部署的全流程开发框架,融合学术与工业经验。
  • 88.38%世界语言为低资源语言,影响超十亿人,需针对性优化。
  • 覆盖技术、语言、文化多维度,适合关注AI公平性的研究者。

多语言大语言模型(MLLMs)是打破语言壁垒、推动人工智能普惠化的重要进展。尽管理论基础已较成熟,但实际落地指南仍零散分散。本文提出一个端到端的完整框架,指导MLLM在生产环境中的开发与部署。主要贡献包括:第一,构建涵盖数据预处理至部署的可操作流程,整合学术研究与工业实践;第二,以Llama2为例,提出提升多语言能力的具体优化策略,包括课程学习平衡高/低资源语言、分词策略及有效采样方法;第三,从技术、语言和文化三重视角进行跨学科分析。研究发现,全球88.38%的语言属于低资源类型,影响超过十亿使用者。通过客服、搜索引擎、机器翻译等真实场景应用,验证了可行解决方案。本综述将理论框架与可落地策略结合,为推动更具包容性和高效性的多语言AI系统提供关键指引。

原文摘要 · Abstract (English)

Multilingual Large Language Models (MLLMs) represent a pivotal advancement in democratizing artificial intelligence across linguistic boundaries. While theoretical foundations are well-established, practical implementation guidelines remain scattered. This work bridges this gap by providing a comprehensive end-to-end framework for developing and deploying MLLMs in production environments. We make three distinctive contributions: First, we present an actionable pipeline from data pre-processing through deployment, integrating insights from academic research and industrial applications. Second, using Llama2 as a case study, we provide detailed optimization strategies for enhancing multilingual capabilities, including curriculum learning approaches for balancing high-resource and low-resource languages, tokenization strategies, and effective sampling methods. Third, we offer an interdisciplinary analysis that considers technical, linguistic, and cultural perspectives in MLLM development. Our findings reveal critical challenges in supporting linguistic diversity, with 88.38% of world languages categorized as low-resource, affecting over a billion speakers. We examine practical solutions through real-world applications in customer service, search engines, and machine translation. By synthesizing theoretical frameworks with production-ready implementation strategies, this survey provides essential guidance for practitioners and researchers working to develop more inclusive and effective multilingual AI systems.

多语言模型AI普惠语言多样性综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。