让多模态模型自己优化生成过程,提升理解与输出的一致性
Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models
- 通过自生成描述词实现模型自我引导的推理机制
- 仅用300样本训练,在多个任务上超越现有方法
- 适合希望提升生成质量的多模态系统开发者
统一多模态模型(UMMs)具备强大的理解能力,但这种能力往往无法有效指导生成。我们将其归因于认知鸿沟:模型缺乏自我优化生成过程的理解。为此,提出内生式重提示(Endogenous Reprompting),将模型理解从被动编码转变为显式的生成推理步骤,即在生成过程中自动生成与目标对齐的描述词。为此设计了SEER(自演化评估器与重提示器)训练框架,基于300个样本的紧凑代理任务「视觉指令扩展」(Visual Instruction Elaboration),采用两阶段内生循环:首先通过可验证奖励的强化学习(RLVR)激活模型潜在评估能力,生成高保真内生奖励信号;其次通过模型奖励思维的强化学习(RLMT)利用该信号优化生成推理策略。实验表明,SEER在评估准确率、重提示效率和生成质量上持续优于先进基线,且不牺牲通用多模态能力。
原文摘要 · Abstract (English)
Unified Multimodal Models (UMMs) exhibit strong understanding, yet this capability often fails to effectively guide generation. We identify this as a Cognitive Gap: the model lacks the understanding of how to enhance its own generation process. To bridge this gap, we propose Endogenous Reprompting, a mechanism that transforms the model's understanding from a passive encoding process into an explicit generative reasoning step by generating self-aligned descriptors during generation. To achieve this, we introduce SEER (Self-Evolving Evaluator and Reprompter), a training framework that establishes a two-stage endogenous loop using only 300 samples from a compact proxy task, Visual Instruction Elaboration. First, Reinforcement Learning with Verifiable Rewards (RLVR) activates the model's latent evaluation ability via curriculum learning, producing a high-fidelity endogenous reward signal. Second, Reinforcement Learning with Model-rewarded Thinking (RLMT) leverages this signal to optimize the generative reasoning policy. Experiments show that SEER consistently outperforms state-of-the-art baselines in evaluation accuracy, reprompting efficiency, and generation quality, without sacrificing general multimodal capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。