arXiv:2501.18648cs.CV2025-01综述被引 19

用多模态大模型增强图像、文本和语音数据,提升模型泛化能力。

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey

  • 利用多模态大模型生成多样化图像、文本和语音数据。
  • 相比传统方法,可有效缓解深度网络过拟合问题。
  • 适合希望提升数据质量与多样性的深度学习研究者。

过去五年间,研究重心从传统机器学习与深度学习转向利用大语言模型(包括多模态)进行数据增强,以提升深度卷积神经网络的泛化能力并缓解过拟合。然而,现有综述多集中于单一模态(如文本或图像)或传统方法,缺乏对最新多模态大模型应用的系统梳理。本文填补这一空白,全面分析近年来基于多模态大模型在图像、文本与音频数据增强中的应用,总结各类方法并讨论当前局限性。同时,结合文献提出潜在解决方案,以提升多模态大模型在数据增强中的有效性。该综述为未来研究提供基础,旨在优化深度学习应用中的数据集质量与多样性。关键词:大模型数据增强、Grok文本增强、DeepSeek图像增强、Grok语音增强、GPT音频增强、语音增强、DeepSeek用于数据增强、DeepSeek R1文本增强、DeepSeek R1图像增强、使用大模型的图像增强、使用大模型的文本增强、大模型数据增强在深度学习中的应用。

原文摘要 · Abstract (English)

In the past five years, research has shifted from traditional Machine Learning (ML) and Deep Learning (DL) approaches to leveraging Large Language Models (LLMs) , including multimodality, for data augmentation to enhance generalization, and combat overfitting in training deep convolutional neural networks. However, while existing surveys predominantly focus on ML and DL techniques or limited modalities (text or images), a gap remains in addressing the latest advancements and multi-modal applications of LLM-based methods. This survey fills that gap by exploring recent literature utilizing multimodal LLMs to augment image, text, and audio data, offering a comprehensive understanding of these processes. We outlined various methods employed in the LLM-based image, text and speech augmentation, and discussed the limitations identified in current approaches. Additionally, we identified potential solutions to these limitations from the literature to enhance the efficacy of data augmentation practices using multimodal LLMs. This survey serves as a foundation for future research, aiming to refine and expand the use of multimodal LLMs in enhancing dataset quality and diversity for deep learning applications. (Surveyed Paper GitHub Repo: https://github.com/WSUAgRobotics/data-aug-multi-modal-llm. Keywords: LLM data augmentation, Grok text data augmentation, DeepSeek image data augmentation, Grok speech data augmentation, GPT audio augmentation, voice augmentation, DeepSeek for data augmentation, DeepSeek R1 text data augmentation, DeepSeek R1 image augmentation, Image Augmentation using LLM, Text Augmentation using LLM, LLM data augmentation for deep learning applications)

多模态大模型数据增强深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。