让多模态模型在生成图像时自我改进,不靠试错,靠内在认知监控。
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
- 用模型自身认知信号指导测试时优化,实现持续学习。
- 在三个主流多模态模型上显著提升组合推理能力,数据有限下仍有效。
- 首次揭示测试时强化学习的关键:认知信号与优化策略的协同作用。
现有统一多模态模型(UMMs)在文本到图像(T2I)生成中的测试时扩展(TTS)方法主要依赖搜索或采样策略,仅实现实例级改进,无法从先前推断中学习或跨相似提示积累知识。为此,我们提出 Meta-TTRL,一种元认知测试时强化学习框架。Meta-TTRL 通过模型内生的元知识监测信号引导测试时参数优化,在测试阶段实现自我改进和能力提升。大量实验表明,Meta-TTRL 在 Janus-Pro-7B、BAGEL 与 Qwen-Image 三个代表性 UMM 上均表现出良好泛化性,在组合推理任务与多个 T2I 基准上取得显著提升,且在数据有限条件下依然有效。我们首次对 UMM 中 T2I 生成的测试时强化学习(TTRL)潜力进行系统分析,揭示有效 TTRL 的关键洞察:元认知协同,即监测信号与模型优化机制对齐,从而实现自我改进。
原文摘要 · Abstract (English)
Existing test-time scaling (TTS) methods for unified multimodal models (UMMs) in text-to-image (T2I) generation primarily rely on search or sampling strategies that produce only instance-level improvements, limiting the ability to learn from prior inferences and accumulate knowledge across similar prompts. To overcome these limitations, we propose Meta-TTRL, a metacognitive test-time reinforcement learning framework. Meta-TTRL performs test-time parameter optimization guided by model-intrinsic monitoring signals derived from the meta-knowledge of UMMs, achieving self-improvement and capability-level improvement at test time. Extensive experiments demonstrate that Meta-TTRL generalizes well across three representative UMMs, including Janus-Pro-7B, BAGEL, and Qwen-Image, achieving significant gains on compositional reasoning tasks and multiple T2I benchmarks with limited data. We provide the first comprehensive analysis to investigate the potential of test-time reinforcement learning (TTRL) for T2I generation in UMMs. Our analysis further reveals a key insight underlying effective TTRL: metacognitive synergy, where monitoring signals align with the model's optimization regime to enable self-improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。