arXiv:2601.10061cs.CVcs.AI2026-01被引 7

用视频模型的逐步推理能力,让文字生成图像更精准。

CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation

  • 通过逐帧推理模拟视觉思考过程,中间帧为可解释的中间步骤。
  • 在GenEval和Imagine-Bench上分别达到0.86和7.468,优于基础视频模型。
  • 专设数据集支持从语义到美学的渐进式生成,减少运动伪影。

近期视频生成模型展现出链式帧(CoF)推理能力,实现逐帧视觉推断。尽管已在迷宫求解、视觉谜题等任务中成功应用,其在文本到图像(T2I)生成中的潜力仍因缺乏明确的视觉推理起点和可解释的中间状态而未被充分挖掘。为此,我们提出CoF-T2I,通过渐进式视觉精炼将CoF推理融入T2I生成,其中中间帧作为显式的推理步骤,最终帧作为输出。为建立这一显式生成流程,我们构建了CoF-Evol-Instruct数据集,记录从语义到美学的生成轨迹。为提升质量并避免运动伪影,每个帧独立编码。实验表明,CoF-T2I显著优于基线视频模型,在GenEval上达0.86,在Imagine-Bench上达7.468,证明视频模型在高质量T2I生成中的巨大潜力。

原文摘要 · Abstract (English)

Recent video generation models have revealed the emergence of Chain-of-Frame (CoF) reasoning, enabling frame-by-frame visual inference. With this capability, video models have been successfully applied to various visual tasks (e.g., maze solving, visual puzzles). However, their potential to enhance text-to-image (T2I) generation remains largely unexplored due to the absence of a clearly defined visual reasoning starting point and interpretable intermediate states in the T2I generation process. To bridge this gap, we propose CoF-T2I, a model that integrates CoF reasoning into T2I generation via progressive visual refinement, where intermediate frames act as explicit reasoning steps and the final frame is taken as output. To establish such an explicit generation process, we curate CoF-Evol-Instruct, a dataset of CoF trajectories that model the generation process from semantics to aesthetics. To further improve quality and avoid motion artifacts, we enable independent encoding operation for each frame. Experiments show that CoF-T2I significantly outperforms the base video model and achieves competitive performance on challenging benchmarks, reaching 0.86 on GenEval and 7.468 on Imagine-Bench. These results indicate the substantial promise of video models for advancing high-quality text-to-image generation.

文本生成图像视频模型视觉推理渐进生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。