arXiv:2609.08282cs.CV2026-09

让多模态模型自己生成图像并自我修正,提升图文生成效果。

Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models

论文配图:Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models
图 1 · 摘自论文原文
  • 用文本提示生成图像,再通过对比反馈优化理解与生成
  • 无需配对图文数据,在多个模型上均提升生成质量
  • 适合研究自进化多模态系统或生成模型的开发者

统一多模态模型将视觉理解与生成整合于单一网络,但两者通常作为独立任务优化。本文提出生成式锚定反馈(GGF),一种仅使用文本提示和模型自身视觉经验的自进化后训练框架。给定一个提示,模型首先生成一幅视觉“梦境”。流级反馈在相同噪声潜在状态中比较文本、图像和修复条件下的预测,将图像锚定的生成方向传递至提示条件。梦境重放锚定通过标注与重想象回放该梦境,训练语义层级证据在重放过程中保持一致,同时分离无关视觉体验。联合优化这两项机制,使生成为理解提供视觉锚定,理解又反向优化后续生成,且无需成对图文监督。在不同理解-生成融合设计的统一模型上实验显示,文本到图像生成性能持续提升,视觉理解略有改善。

原文摘要 · Abstract (English)

Unified multimodal models integrate visual understanding and generation within a single network, yet the two capabilities are commonly optimized as separate tasks. We introduce Generative Grounding Feedback(GGF), a self-evolving post-training framework that uses only text prompts and the model's own visual experience. Given a prompt, the model first generates a visual ``dream.'' Flow-level feedback compares text-, image-, and repair-conditioned predictions at the same noisy latent state, transferring image-grounded generation directions to the prompt condition. Dream replay grounding replays this dream through captioning and re-imagination, training claim-level evidence to remain consistent across the replay while separating unrelated visual experiences. Jointly optimized, these two directions let generation provide visual grounding for understanding and understanding refine subsequent generation without paired image--text supervision. Experiments across unified models with different understanding--generation integration designs show consistent improvements in text-to-image generation together with modest gains in visual understanding.

多模态生成模型自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。