让多模态模型自己教自己生成,提升图像理解与创作能力。
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
- 拆分模型为提议者、求解者和裁判,通过自博弈生成高质量训练信号。
- 在六个图像生成基准上显著超越基线,最高提升6.5分,部分达新纪录。
- 无需外部数据,适合追求自监督生成的AI研究者和开发者。
统一多模态模型(UMMs)在跨模态理解方面取得显著进展,但在利用内部知识进行高质量生成方面仍存在明显差距。我们将其称为‘传导性失语’——模型能准确理解多模态输入,却难以生成忠实且可控的内容。为此,我们提出UniCorn,一个无需外部数据或教师监督的自改进框架。通过将单一UMM划分为提议者、求解者和裁判三个协作角色,实现自博弈生成,并采用认知模式重建技术,将隐含理解提炼为显式生成信号。为验证多模态一致性恢复,我们引入UniCycle,一种基于文本→图像→文本重构循环的周期一致性基准。大量实验表明,UniCorn在六个通用图像生成基准上全面超越基线模型,尤其在TIIF(73.8)、DPG(86.8)、CompBench(88.5)和UniCycle上达到当前最佳性能,同时在WISE上提升+5.0,在OneIG上提升+6.5。结果表明,该方法显著增强了文本到图像生成能力,且保持强理解力,证明了完全自监督优化在统一多模态智能中的可扩展性。
原文摘要 · Abstract (English)
While Unified Multimodal Models (UMMs) have achieved remarkable success in cross-modal comprehension, a significant gap persists in their ability to leverage such internal knowledge for high-quality generation. We formalize this discrepancy as Conduction Aphasia, a phenomenon where models accurately interpret multimodal inputs but struggle to translate that understanding into faithful and controllable synthesis. To address this, we propose UniCorn, a simple yet elegant self-improvement framework that eliminates the need for external data or teacher supervision. By partitioning a single UMM into three collaborative roles: Proposer, Solver, and Judge, UniCorn generates high-quality interactions via self-play and employs cognitive pattern reconstruction to distill latent understanding into explicit generative signals. To validate the restoration of multimodal coherence, we introduce UniCycle, a cycle-consistency benchmark based on a Text to Image to Text reconstruction loop. Extensive experiments demonstrate that UniCorn achieves comprehensive and substantial improvements over the base model across six general image generation benchmarks. Notably, it achieves SOTA performance on TIIF(73.8), DPG(86.8), CompBench(88.5), and UniCycle while further delivering substantial gains of +5.0 on WISE and +6.5 on OneIG. These results highlight that our method significantly enhances T2I generation while maintaining robust comprehension, demonstrating the scalability of fully self-supervised refinement for unified multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。