构建了600万对高质量图文数据集,助力文本生成图像模型精准微调。
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning
- 融合合成与真实图像,严格筛选确保图文一致性和画质。
- 含超600万图文对,数据量接近预训练规模,支持高精度微调。
- 适用于追求生成质量与指令遵循的科研及工业级T2I研发者。
高质量且开源的数据集仍是文本到图像(T2I)微调的主要瓶颈。尽管模型架构和训练流程快速进步,多数公开微调数据集仍存在分辨率低、图文对齐差或多样性不足的问题,导致开源研究模型与企业级模型间存在明显性能差距。本文提出Fine-T2I,一个大规模、高质量、完全开源的T2I微调数据集。Fine-T2I涵盖10种任务组合、32个提示类别、11种视觉风格和5种提示模板,结合强模型生成的合成图像与专业摄影师提供的真实图像。所有样本均经过严格过滤,确保图文对齐、视觉保真度和提示质量,初始候选样本中超过95%被剔除。最终数据集包含超过600万对文本-图像数据,约2TB存储空间,规模接近预训练数据集,同时保持微调级别质量。在多种预训练扩散模型和自回归模型上,基于Fine-T2I进行微调,显著提升生成质量与指令遵循能力,经人工评估、视觉对比和自动指标验证。本研究以开放许可发布Fine-T2I,旨在缩小开源社区中T2I微调的数据差距。
原文摘要 · Abstract (English)
High-quality and open datasets remain a major bottleneck for text-to-image (T2I) fine-tuning. Despite rapid progress in model architectures and training pipelines, most publicly available fine-tuning datasets suffer from low resolution, poor text-image alignment, or limited diversity, resulting in a clear performance gap between open research models and enterprise-grade models. In this work, we present Fine-T2I, a large-scale, high-quality, and fully open dataset for T2I fine-tuning. Fine-T2I spans 10 task combinations, 32 prompt categories, 11 visual styles, and 5 prompt templates, and combines synthetic images generated by strong modern models with carefully curated real images from professional photographers. All samples are rigorously filtered for text-image alignment, visual fidelity, and prompt quality, with over 95% of initial candidates removed. The final dataset contains over 6 million text-image pairs, around 2 TB on disk, approaching the scale of pretraining datasets while maintaining fine-tuning-level quality. Across a diverse set of pretrained diffusion and autoregressive models, fine-tuning on Fine-T2I consistently improves both generation quality and instruction adherence, as validated by human evaluation, visual comparison, and automatic metrics. We release Fine-T2I under an open license to help close the data gap in T2I fine-tuning in the open community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。