构建首个图文生成持续训练基准,评估模型长期学习能力。
T2I-ConBench: Text-to-Image Benchmark for Continual Post-training

- 设计覆盖定制与领域增强的双场景评测框架。
- 发现现有方法均无法在所有维度上表现优异。
- 开源数据集与工具,推动持续训练研究进展。
持续后训练使单个文生图扩散模型可低成本学习新任务,但简单训练会导致预训练知识遗忘并破坏零样本组合能力。我们发现缺乏标准化评估协议制约了该方向研究。为此,提出 T2I-ConBench,一个面向文生图模型持续后训练的统一基准。该基准聚焦物品定制与领域增强两类实际场景,从泛化保留、目标任务性能、灾难性遗忘和跨任务泛化四个维度进行评估,结合自动化指标、人类偏好建模与视觉语言问答实现全面测评。我们在三个真实任务序列上对十种代表性方法进行评测,发现无一方法在所有维度上均表现最优,即使联合‘理想’训练也无法在每项任务中成功,跨任务泛化问题仍未解决。我们开源全部数据集、代码与评估工具,以加速文生图持续后训练研究。
原文摘要 · Abstract (English)
Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of pretrained knowledge and undermines zero-shot compositionality. We observe that the absence of a standardized evaluation protocol hampers related research for continual post-training. To address this, we introduce T2I-ConBench, a unified benchmark for continual post-training of text-to-image models. T2I-ConBench focuses on two practical scenarios, item customization and domain enhancement, and analyzes four dimensions: (1) retention of generality, (2) target-task performance, (3) catastrophic forgetting, and (4) cross-task generalization. It combines automated metrics, human-preference modeling, and vision-language QA for comprehensive assessment. We benchmark ten representative methods across three realistic task sequences and find that no approach excels on all fronts. Even joint "oracle" training does not succeed for every task, and cross-task generalization remains unsolved. We release all datasets, code, and evaluation tools to accelerate research in continual post-training for text-to-image models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。