arXiv:2507.08648cs.CVcs.AI2025-07被引 6

用多智能体自动构建高质量真实图像数据集

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images

  • 四智能体协同,结合多模态大模型与图像优化工具
  • 可扩展现有数据集或从零创建新数据集,支持多种视觉任务训练
  • 适合需要高效构建真实世界图像数据集的研究者

构建图像数据集通常依赖耗时低效的人工采集与标注。大模型虽可生成数据,但真实世界数据仍更具价值。为此,本文提出DatasetAgent——一种基于多智能体协作系统的自动化真实图像数据集构建方法。该系统由四个配备多模态大语言模型(MLLMs)的智能体及一套图像优化工具包组成,可根据用户指定需求生成高质量图像数据集。通过两类实验验证:在多个开源数据集上进行数据集扩展与全新构建。所生成的数据集用于训练图像分类、目标检测和图像分割等多种视觉模型,证明其有效性。

原文摘要 · Abstract (English)

Common knowledge indicates that the process of constructing image datasets usually depends on the time-intensive and inefficient method of manual collection and annotation. Large models offer a solution via data generation. Nonetheless, real-world data are obviously more valuable comparing to artificially intelligence generated data, particularly in constructing image datasets. For this reason, we propose a novel method for auto-constructing datasets from real-world images by a multiagent collaborative system, named as DatasetAgent. By coordinating four different agents equipped with Multi-modal Large Language Models (MLLMs), as well as a tool package for image optimization, DatasetAgent is able to construct high-quality image datasets according to user-specified requirements. In particular, two types of experiments are conducted, including expanding existing datasets and creating new ones from scratch, on a variety of open-source datasets. In both cases, multiple image datasets constructed by DatasetAgent are used to train various vision models for image classification, object detection, and image segmentation.

数据集构建多智能体图像生成MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。