arXiv:2412.02142cs.CVcs.AI2024-12综述被引 21

系统梳理个性化多模态大模型的架构、训练与应用方法

Personalized Multimodal Large Language Models: A Survey

  • 提出直观分类框架,归纳用户个性化技术
  • 总结主流任务与评估指标,梳理关键数据集
  • 适合关注多模态模型定制化的研究者阅读

多模态大语言模型(MLLMs)因在文本、图像、音频等多模态融合任务中表现卓越而日益重要。本文全面综述个性化多模态大语言模型,重点分析其架构设计、训练方法与应用场景。提出一种直观的分类体系,用于归类个性化技术,并讨论其结合与适配策略,阐明各自优势与原理。总结现有研究中的个性化任务及常用评估指标,梳理可用于基准测试的数据集。最后,指出当前亟待解决的关键挑战。本综述旨在为研究人员和实践者理解并推进个性化多模态大模型的发展提供参考。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have become increasingly important due to their state-of-the-art performance and ability to integrate multiple data modalities, such as text, images, and audio, to perform complex tasks with high accuracy. This paper presents a comprehensive survey on personalized multimodal large language models, focusing on their architecture, training methods, and applications. We propose an intuitive taxonomy for categorizing the techniques used to personalize MLLMs to individual users, and discuss the techniques accordingly. Furthermore, we discuss how such techniques can be combined or adapted when appropriate, highlighting their advantages and underlying rationale. We also provide a succinct summary of personalization tasks investigated in existing research, along with the evaluation metrics commonly used. Additionally, we summarize the datasets that are useful for benchmarking personalized MLLMs. Finally, we outline critical open challenges. This survey aims to serve as a valuable resource for researchers and practitioners seeking to understand and advance the development of personalized multimodal large language models.

多模态大模型个性化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。