首份综述系统梳理多模态大模型自提升方法与挑战
Self-Improvement in Multimodal Large Language Models: A Survey
- 从数据收集、组织到模型优化三方面梳理自提升技术路径
- 总结多模态大模型自提升的评估方法与实际应用案例
- 适合关注多模态AI自我进化方向的研究者与工程师
近年来,大语言模型的自提升技术在不显著增加成本(尤其是人力成本)的情况下有效提升了模型能力。尽管该领域仍处于早期阶段,但其向多模态领域的拓展具有巨大潜力,可利用多样化数据源并发展更通用的自提升模型。本文是首篇系统综述多模态大语言模型(MLLMs)自提升的研究。我们从数据收集、数据组织和模型优化三个角度对现有文献进行了结构化梳理,以促进该领域进一步发展。同时涵盖了常用评估方法与下游应用场景。最后,总结了当前开放挑战与未来研究方向。
原文摘要 · Abstract (English)
Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potential for leveraging diverse data sources and developing more general self-improving models. This survey is the first to provide a comprehensive overview of self-improvement in Multimodal LLMs (MLLMs). We provide a structured overview of the current literature and discuss methods from three perspectives: 1) data collection, 2) data organization, and 3) model optimization, to facilitate the further development of self-improvement in MLLMs. We also include commonly used evaluations and downstream applications. Finally, we conclude by outlining open challenges and future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。