arXiv:2505.15172cs.CV2025-05被引 1

用新指标选细描述,小数据也能生成高质量图像。

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation

  • 基于图像覆盖度和物体描述精细度构建细节度评估指标
  • 仅用20%优质数据训练,效果超越全量数据和按长度选数方法
  • 适合追求数据效率与生成质量的图像生成研究者

使用详细描述训练文本到图像(T2I)模型可显著提升生成质量。现有方法多依赖字数等简单指标衡量描述细节度。本文提出新指标:图像覆盖率(ICR)衡量描述是否涵盖图像所有区域/物体,平均物体细节度(AOD)量化每个物体描述的细致程度。在使用ShareGPT4V标注的COCO数据集上实验表明,基于高ICR与高AOD的描述训练的T2I模型,在DPG及其他基准测试中表现更优。值得注意的是,仅用20%精选数据训练即超越全量数据训练及长度筛选法,显著提升对齐与重建能力。结果凸显细节感知指标在T2I数据筛选中的关键作用,优于传统长度启发式方法。

原文摘要 · Abstract (English)

Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length to represent the detailness of the caption in the T2I training set. In this paper, we propose a new metric to estimate caption detailness based on two aspects: image coverage rate (ICR), which evaluates whether the caption covers all regions/objects in the image, and average object detailness (AOD), which quantifies the detailness of each object's description. Through experiments on the COCO dataset using ShareGPT4V captions, we demonstrate that T2I models trained on high-ICR and -AOD captions achieve superior performance on DPG and other benchmarks. Notably, our metric enables more effective data selection-training on only 20% of full data surpasses both full-dataset training and length-based selection method, improving alignment and reconstruction ability. These findings highlight the critical role of detail-aware metrics over length-based heuristics in caption selection for T2I tasks.

文本生成图像数据效率描述质量图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。