arXiv:2605.16961cs.CVcs.AI2026-05

让模型的推理直接控制图像生成,提升复杂场景理解与生成一致性。

Latent Action Control for Reasoning-Guided Unified Image Generation

论文配图:Latent Action Control for Reasoning-Guided Unified Image Generation
图 1 · 摘自论文原文
  • 将推理过程转化为隐式连续动作,嵌入统一生成器中
  • 在多个评测集上显著提升空间关系与知识依赖型生成效果
  • 适合需要精准语义控制的图像生成任务

统一多模态模型虽能共享主干网络实现视觉理解与图像生成,但理解信息未必能转化为可控生成。本文提出隐式动作控制(Latent Action Control, LAC),将推理过程表示为统一生成器内部的连续隐动作,通过角色结构化的潜在轨迹进行规划、内视草图、诊断与优化,并将这些动作注入条件流模型中,无需生成中间图像或推理标记。由于动作轨迹不可观测,LAC通过仅训练阶段渲染的语义先验、草图特征与监督终止信号,结合变分潜动作对齐学习,并使用潜流GRPO对齐潜动作到图像的演化过程。该方法打通了从推断关系、绑定和知识线索到生成的控制路径。在BAGEL-7B-MoT上,LAC在GenEval、WISE和T2I-CompBench上均显著提升组合与知识引导生成性能,尤其在空间关系、属性绑定及世界知识敏感提示下增益最大。消融实验与潜空间干预表明,学习到的动作轨迹被生成器实际消耗,证明理解只有在生成时可行动才真正助力统一生成。

原文摘要 · Abstract (English)

Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may infer objects, relations, or knowledge cues but fail to instantiate them in the generated image. We propose Latent Action Control (LAC), which makes reasoning actionable by representing it as hidden continuous actions inside a unified generator. Given a prompt, LAC rolls out a role-structured latent trajectory for planning, internal visual drafting, diagnosis, and refinement, and injects these actions into the hidden stream that conditions flow-based generation, without producing reasoning tokens or intermediate images. Since such action trajectories are unobserved, LAC learns them through prior-guided variational latent action alignment from training-only rendered semantic priors, draft image features, and supervised halting signals, followed by Latent-Flow GRPO to align the latent-to-image rollout with terminal visual feedback. This provides a control path from inferred relations, bindings, and knowledge cues to the generation process. Instantiated on BAGEL-7B-MoT, LAC consistently improves compositional and knowledge-grounded generation across GenEval, WISE, and T2I-CompBench, with the largest gains on spatial relations, attribute binding, and world-knowledge-sensitive prompts. Ablations and latent interventions show that the learned action trajectory is consumed by the generator, suggesting that unified generation benefits when understanding is not only encoded, but made actionable during generation.

图像生成推理控制统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。