arXiv:2509.17458cs.CVcs.CL2025-09中稿 · TMLR被引 2

通过智能选择奖励信号,让图像生成更准确匹配复杂提示。

CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration

  • 根据人类判断相关性选择奖励信号,统一优化与探索策略。
  • 在两个评测集上平均对齐得分提升16%和11%,优于现有方法。
  • 适合需要精准描述复杂场景的图像生成应用,如设计、创意工作。

文本到图像扩散模型(如 Stable Diffusion)虽能生成高质量且多样化的图像,但在描述复杂物体关系、属性或空间布局时常出现组合对齐失败。近期推理阶段方法通过在无模型微调前提下,利用评分函数引导初始噪声优化或探索来改善对齐。然而,单独使用优化易因初始化不佳或搜索轨迹不利而停滞;探索则需大量样本才能找到满意输出。我们分析发现,单一奖励指标或随意组合均无法可靠捕捉所有组合性特征,导致引导效果弱且不一致。为此,提出统一框架 CARINOX,结合噪声优化与探索,并基于与人类判断的相关性进行奖励选择。在覆盖多种组合挑战的两个互补基准测试中,CARINOX 在 T2I-CompBench++ 上平均对齐得分提升 +16%,在 HRS 基准上提升 +11%,在所有主要类别中持续优于当前最优的优化与探索方法,同时保持图像质量和多样性。

原文摘要 · Abstract (English)

Text-to-image diffusion models, such as Stable Diffusion, can produce high-quality and diverse images but often fail to achieve compositional alignment, particularly when prompts describe complex object relationships, attributes, or spatial arrangements. Recent inference-time approaches address this by optimizing or exploring the initial noise under the guidance of reward functions that score text-image alignment without requiring model fine-tuning. While promising, each strategy has intrinsic limitations when used alone: optimization can stall due to poor initialization or unfavorable search trajectories, whereas exploration may require a prohibitively large number of samples to locate a satisfactory output. Our analysis further shows that neither single reward metrics nor ad-hoc combinations reliably capture all aspects of compositionality, leading to weak or inconsistent guidance. To overcome these challenges, we present Category-Aware Reward-based Initial Noise Optimization and Exploration (CARINOX), a unified framework that combines noise optimization and exploration with a principled reward selection procedure grounded in correlation with human judgments. Evaluations on two complementary benchmarks covering diverse compositional challenges show that CARINOX raises average alignment scores by +16% on T2I-CompBench++ and +11% on the HRS benchmark, consistently outperforming state-of-the-art optimization and exploration-based methods across all major categories, while preserving image quality and diversity. The project page is available at https://amirkasaei.com/carinox/.

图像生成扩散模型提示对齐推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。