arXiv:2509.22761cs.CVcs.AI2025-09被引 7

测试时联合推理图文,提升多模态图像生成质量

MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning

  • 在统一潜在空间中同步优化图文表示,测试时动态推理
  • 在WISE数据集上得分0.63,较基线提升80%以上
  • 适合需要精细跨模态理解的图像生成任务

推理增强的机器学习系统在多个领域表现出色,包括图像生成。然而,现有基于推理的图像生成方法要么仅限于单模态(图像或文本)推理,要么依赖高质量推理数据进行微调。为解决这些局限,我们提出MILR,一种测试时方法,在统一潜在向量空间中联合推理图像和文本。推理通过搜索离散图像与文本标记的向量表示实现,实际采用策略梯度方法,由图像质量评判器引导。我们在支持语言推理的统一多模态理解与生成(MUG)框架中实例化MILR,该框架天然支持图像合成前的语言推理,从而促进跨模态推理。中间模型输出作为需优化的统一潜在空间,使MILR完全在测试时运行。我们在GenEval、T2I-CompBench和WISE上评估MILR,所有基准均达当前最优结果。尤其在知识密集型WISE上,MILR总体得分为0.63,相比基线提升80%。进一步分析表明,统一潜在空间中的联合推理是其高性能的关键。定性研究还揭示MILR具备非平凡的时间与文化推理能力,验证了该推理方法的有效性。

原文摘要 · Abstract (English)

Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image generation either restrict reasoning to a single modality (image or text) or rely on high-quality reasoning data for fine-tuning. To tackle these limitations, we propose MILR, a test-time method that jointly reasons over image and text in a unified latent vector space. Reasoning in MILR is performed by searching through vector representations of discrete image and text tokens. Practically, this is implemented via the policy gradient method, guided by an image quality critic. We instantiate MILR within the unified multimodal understanding and generation (MUG) framework that natively supports language reasoning before image synthesis and thus facilitates cross-modal reasoning. The intermediate model outputs, which are to be optimized, serve as the unified latent space, enabling MILR to operate entirely at test time. We evaluate MILR on GenEval, T2I-CompBench, and WISE, achieving state-of-the-art results on all benchmarks. Notably, on knowledge-intensive WISE, MILR attains an overall score of 0.63, improving over the baseline by 80%. Our further analysis indicates that joint reasoning in the unified latent space is the key to its strong performance. Moreover, our qualitative studies reveal MILR's non-trivial ability in temporal and cultural reasoning, highlighting the efficacy of our reasoning method.

多模态生成测试时推理图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。