构建1240万张无版权图像文本数据集,支持AI绘画模型训练
Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms
- 收集1240万张公共领域图像,自动生成合成图文对
- 数据量达1240万,是目前最大公开图像文本数据集
- 通过社区治理平台保障数据安全与长期可复现性
我们提出Public Domain 12M(PD12M),一个包含1240万张高质量公共领域及CC0许可图像的图文数据集,用于训练文本到图像模型。该数据集是迄今最大的公共领域图像文本数据集,规模足以训练基础模型,同时最大限度降低版权风险。通过Source.Plus平台,我们引入新型社区驱动的数据治理机制,有效减少潜在危害并支持长期可复现性。
原文摘要 · Abstract (English)
We present Public Domain 12M (PD12M), a dataset of 12.4 million high-quality public domain and CC0-licensed images with synthetic captions, designed for training text-to-image models. PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。