arXiv:2506.05673cs.LGcs.AI2025-06

构建高质量图像数据集,提升视觉模型训练效果

Peer-Ranked Precision: Creating a Foundational Dataset for Fine-Tuning Vision Models from DataSeeds' Annotated Imagery

  • 基于人类评分构建10610张高质图像数据集
  • 在基准测试中显著提升模型性能
  • 适合用于商业与多模态AI训练

现代人工智能模型,尤其是基于扩散的计算机视觉与图像生成模型,正从传统的“模型中心”方法转向更注重数据质量的“数据中心”范式。为此,我们推出DataSeeds.AI样本数据集(简称DSD),包含约10,610张经人类同行评分的高质量摄影图像,并附有详尽的多层级标注。DSD仅占DataSeeds.AI超1亿图像目录的一小部分,但为商业及多模态AI开发提供了可扩展的基础。通过深入分析,我们在特定模型上验证了DSD带来的量化性能提升,并公开了评估所用代码与训练模型。

原文摘要 · Abstract (English)

The development of modern Artificial Intelligence (AI) models, particularly diffusion-based models employed in computer vision and image generation tasks, is undergoing a paradigmatic shift in development methodologies. Traditionally dominated by a "Model Centric" approach, in which performance gains were primarily pursued through increasingly complex model architectures and hyperparameter optimization, the field is now recognizing a more nuanced "Data-Centric" approach. This emergent framework foregrounds the quality, structure, and relevance of training data as the principal driver of model performance. To operationalize this paradigm shift, we introduce the DataSeeds.AI sample dataset (the "DSD"), initially comprised of approximately 10,610 high-quality human peer-ranked photography images accompanied by extensive multi-tier annotations. The DSD is a foundational computer vision dataset designed to usher in a new standard for commercial image datasets. Representing a small fraction of DataSeeds.AI's 100 million-plus image catalog, the DSD provides a scalable foundation necessary for robust commercial and multimodal AI development. Through this in-depth exploratory analysis, we document the quantitative improvements generated by the DSD on specific models against known benchmarks and make the code and the trained models used in our evaluation publicly available.

数据集视觉模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。