arXiv:2603.26866cs.CVcs.AI2026-03

不用删数据,用质量标签训练模型,生成效果更好

LACON: Training Text-to-Image Model from Uncurated Data

  • 把图像质量分数当标签,让模型学会分辨好坏数据
  • 相同算力下,生成质量超越只用筛选数据的基线方法
  • 适合想用海量杂乱数据训练高质量生成模型的研究者

现代文本到图像生成的成功很大程度上依赖于大规模高质量数据集。目前这些数据集采用先过滤后训练的范式,基于低质量数据有害于模型性能的假设,主动剔除大量原始数据。被丢弃的低质量数据是否真无用?本文对此提出质疑。我们提出 LACON(Labeling-and-Conditioning)框架,利用未清理数据的内在分布。不进行过滤,而是将美学评分、水印概率等质量信号转化为显式的量化条件标签,使生成模型学习从劣质到优质数据的完整谱系。通过显式学习高低质量内容的边界,LACON 在相同计算预算下,生成质量优于仅在筛选数据上训练的基线模型,证明了未清理数据的巨大价值。

原文摘要 · Abstract (English)

The success of modern text-to-image generation is largely attributed to massive, high-quality datasets. Currently, these datasets are curated through a filter-first paradigm that aggressively discards low-quality raw data based on the assumption that it is detrimental to model performance. Is the discarded bad data truly useless, or does it hold untapped potential? In this work, we critically re-examine this question. We propose LACON (Labeling-and-Conditioning), a novel training framework that exploits the underlying uncurated data distribution. Instead of filtering, LACON re-purposes quality signals, such as aesthetic scores and watermark probabilities as explicit, quantitative condition labels. The generative model is then trained to learn the full spectrum of data quality, from bad to good. By learning the explicit boundary between high- and low-quality content, LACON achieves superior generation quality compared to baselines trained only on filtered data using the same compute budget, proving the significant value of uncurated data.

文本生成数据利用生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。