arXiv:2503.16585cs.CLcs.CV2025-03综述被引 14

综述分布式大模型与多模态大模型的进展与挑战

Distributed LLMs and Multimodal Large Language Models: A Survey on Advances, Challenges, and Future Directions

  • 系统梳理分布式训练、推理等全链路技术
  • 涵盖文本、图像、音频多模态大模型应用
  • 适合关注大模型落地与隐私计算的研究者

语言模型(LM)是基于大规模数据集(如文本)估算词序列概率以预测语言模式的机器学习模型,在自然语言处理任务中广泛应用,包括自动补全和机器翻译。尽管更大规模的数据通常能提升模型性能,但受限于算力和资源,可扩展性仍是挑战。分布式计算策略为提升可扩展性和应对日益增长的计算需求提供了关键解决方案。此外,训练和部署中使用敏感数据引发重大隐私问题。近期研究聚焦于发展去中心化技术,实现分布式训练与推理,利用多样化计算资源并支持边缘AI。本文综述了各类语言模型(包括大语言模型LLMs、视觉语言模型VLMs、多模态大语言模型MLLMs和小语言模型SLMs)的分布式解决方案。虽然LLMs专注于文本处理与生成,而MLLMs设计用于处理多种模态数据(如文本、图像、音频)并进行融合以拓展应用场景。为此,本文回顾了多模态大模型全链路的关键进展,包括分布式训练、推理、微调与部署,并识别了现有贡献、局限及未来改进方向。进一步地,按六个主要去中心化焦点对文献进行分类。我们的分析指出了当前实现分布式语言模型方法中的空白,并提出了未来研究方向,强调需开发新方案以增强分布式语言模型的鲁棒性与适用性。

原文摘要 · Abstract (English)

Language models (LMs) are machine learning models designed to predict linguistic patterns by estimating the probability of word sequences based on large-scale datasets, such as text. LMs have a wide range of applications in natural language processing (NLP) tasks, including autocomplete and machine translation. Although larger datasets typically enhance LM performance, scalability remains a challenge due to constraints in computational power and resources. Distributed computing strategies offer essential solutions for improving scalability and managing the growing computational demand. Further, the use of sensitive datasets in training and deployment raises significant privacy concerns. Recent research has focused on developing decentralized techniques to enable distributed training and inference while utilizing diverse computational resources and enabling edge AI. This paper presents a survey on distributed solutions for various LMs, including large language models (LLMs), vision language models (VLMs), multimodal LLMs (MLLMs), and small language models (SLMs). While LLMs focus on processing and generating text, MLLMs are designed to handle multiple modalities of data (e.g., text, images, and audio) and to integrate them for broader applications. To this end, this paper reviews key advancements across the MLLM pipeline, including distributed training, inference, fine-tuning, and deployment, while also identifying the contributions, limitations, and future areas of improvement. Further, it categorizes the literature based on six primary focus areas of decentralization. Our analysis describes gaps in current methodologies for enabling distributed solutions for LMs and outline future research directions, emphasizing the need for novel solutions to enhance the robustness and applicability of distributed LMs.

大模型分布式多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。