让图像生成的思维链更短更准,提速54%还不丢质量。
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
- 用自适应奖励机制引导生成更简练的思维链
- 推理长度减少54%,多基准测试质量不降反升
- 适合追求高效图像生成的开发者和研究者
自回归多模态大模型在图像生成中日益流行,新方法采用思维链(CoT)推理,在生成图像前扩展用户输入以提升对齐度与细节。然而,该策略常引发冗余——我们称之为视觉过度思考,导致计算成本上升,并可能引入与原提示矛盾的细节。本文提出ShortCoTI,一种轻量级优化框架,通过自适应奖励函数鼓励更简洁的思维链,奖励强度随任务难度动态调整。将其融入强化学习范式后,推理长度平均减少54%,在T2I-CompBench与GenEval多个基准上保持或略微提升质量。定性分析显示,该方法消除冗长解释与重复修正,生成的推理提示更简洁且语义丰富。ShortCoTI在不牺牲图像保真度与视觉美感的前提下,显著提升生成效率。
原文摘要 · Abstract (English)
Autoregressive multimodal large language models have recently gained popularity for image generation, driven by advances in foundation models. To enhance alignment and detail, newer approaches employ chain-of-thought (CoT) reasoning, expanding user inputs into elaborated prompts prior to image synthesis. However, this strategy can introduce unnecessary redundancy -- a phenomenon we call visual overthinking -- which increases computational costs and can introduce details that contradict the original prompt. In this work, we explore how to generate more concise CoT sequences for more efficient image generation. We introduce ShortCoTI, a lightweight optimization framework that encourages more concise CoT while preserving output image quality. ShortCoTI rewards more concise prompts with an adaptive function that scales according to an estimated difficulty for each task. Incorporating this reward into a reinforcement learning paradigm reduces prompt reasoning length by 54% while maintaining or slightly improving quality metrics across multiple benchmarks (T2I-CompBench, GenEval). Qualitative analysis shows that our method eliminates verbose explanations and repetitive refinements, producing reasoning prompts that are both concise and semantically rich. As a result, ShortCoTI improves computational efficiency without compromising the fidelity or visual appeal of generated images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。