arXiv:2512.22374cs.CVcs.AI2025-12被引 2

自评估机制让文生图模型任意步数生成,快且质量高。

Self-Evaluation Unlocks Any-Step Text-to-Image Generation

  • 自评机制让模型自我监督,无需预训练教师
  • 10步内生成质量媲美50步传统模型
  • 统一模型支持极速生成与高质量长轨迹

我们提出自评估模型(Self-E),一种从零开始训练的文生图新方法,支持任意步数推理。Self-E 的训练方式类似流匹配模型,同时引入新颖的自评估机制:利用当前得分估计自我评估生成样本,充当动态自教师。不同于依赖局部监督的传统扩散或流模型(需多步推理),也不同于需要预训练教师的蒸馏方法,该机制结合即时局部学习与自驱动全局匹配,使模型从零训练即可达到高质量,尤其在极低步数下表现优异。大规模文生图基准测试表明,Self-E 在少步生成上表现卓越,且在50步时可与顶尖流匹配模型比肩。进一步发现其性能随推理步数单调提升,单模型即可实现超快速生成与高质量长轨迹采样。据我们所知,Self-E 是首个从零训练、支持任意步数的文生图模型,提供高效可扩展的统一生成框架。

原文摘要 · Abstract (English)

We introduce the Self-Evaluating Model (Self-E), a novel, from-scratch training approach for text-to-image generation that supports any-step inference. Self-E learns from data similarly to a Flow Matching model, while simultaneously employing a novel self-evaluation mechanism: it evaluates its own generated samples using its current score estimates, effectively serving as a dynamic self-teacher. Unlike traditional diffusion or flow models, it does not rely solely on local supervision, which typically necessitates many inference steps. Unlike distillation-based approaches, it does not require a pretrained teacher. This combination of instantaneous local learning and self-driven global matching bridges the gap between the two paradigms, enabling the training of a high-quality text-to-image model from scratch that excels even at very low step counts. Extensive experiments on large-scale text-to-image benchmarks show that Self-E not only excels in few-step generation, but is also competitive with state-of-the-art Flow Matching models at 50 steps. We further find that its performance improves monotonically as inference steps increase, enabling both ultra-fast few-step generation and high-quality long-trajectory sampling within a single unified model. To our knowledge, Self-E is the first from-scratch, any-step text-to-image model, offering a unified framework for efficient and scalable generation.

文生图任意步数自评估扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。