arXiv:2412.17759cs.AIcs.CV2024-12综述被引 23

梳理主流多模态数据集,助你快速了解该领域研究现状。

Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy

  • 系统整理多模态模型训练与应用所需的数据集
  • 归纳涵盖文本、图像、音视频的典型数据集类别
  • 适合想入门或评估多模态模型的研究者参考

多模态学习是人工智能中迅速发展的领域,旨在通过整合分析文本、图像、音频和视频等多种数据类型,构建更通用、更鲁棒的系统。受人类多感官信息融合能力启发,该方法支持文本生成视频、视觉问答、图像描述等应用。本文综述了近年来支撑多模态语言模型(MLLMs)的大规模多模态数据集发展。大规模多模态数据集对模型训练与全面测试至关重要。研究重点分析了用于训练、特定任务和实际应用的数据集,并强调基准数据集在评估模型性能、可扩展性与适用性方面的重要作用。随着多模态学习不断演进,克服这些挑战将推动人工智能研究与应用迈向新高度。

原文摘要 · Abstract (English)

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the human ability to assimilate information through many senses, this method enables applications such as text-to-video conversion, visual question answering, and image captioning. Recent developments in datasets that support multimodal language models (MLLMs) are highlighted in this overview. Large-scale multimodal datasets are essential because they allow for thorough testing and training of these models. With an emphasis on their contributions to the discipline, the study examines a variety of datasets, including those for training, domain-specific tasks, and real-world applications. It also emphasizes how crucial benchmark datasets are for assessing models' performance in a range of scenarios, scalability, and applicability. Since multimodal learning is always changing, overcoming these obstacles will help AI research and applications reach new heights.

多模态数据集综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。