arXiv:2601.20305cs.AI2026-01被引 4

让多模态模型自己优化生成过程,提升理解与输出的一致性

Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models

  • 通过自生成描述词实现模型自我引导的推理机制
  • 仅用300样本训练,在多个任务上超越现有方法
  • 适合希望提升生成质量的多模态系统开发者

统一多模态模型(UMMs)具备强大的理解能力,但这种能力往往无法有效指导生成。我们将其归因于认知鸿沟:模型缺乏自我优化生成过程的理解。为此,提出内生式重提示(Endogenous Reprompting),将模型理解从被动编码转变为显式的生成推理步骤,即在生成过程中自动生成与目标对齐的描述词。为此设计了SEER(自演化评估器与重提示器)训练框架,基于300个样本的紧凑代理任务「视觉指令扩展」(Visual Instruction Elaboration),采用两阶段内生循环:首先通过可验证奖励的强化学习(RLVR)激活模型潜在评估能力,生成高保真内生奖励信号;其次通过模型奖励思维的强化学习(RLMT)利用该信号优化生成推理策略。实验表明,SEER在评估准确率、重提示效率和生成质量上持续优于先进基线,且不牺牲通用多模态能力。

原文摘要 · Abstract (English)

Unified Multimodal Models (UMMs) exhibit strong understanding, yet this capability often fails to effectively guide generation. We identify this as a Cognitive Gap: the model lacks the understanding of how to enhance its own generation process. To bridge this gap, we propose Endogenous Reprompting, a mechanism that transforms the model's understanding from a passive encoding process into an explicit generative reasoning step by generating self-aligned descriptors during generation. To achieve this, we introduce SEER (Self-Evolving Evaluator and Reprompter), a training framework that establishes a two-stage endogenous loop using only 300 samples from a compact proxy task, Visual Instruction Elaboration. First, Reinforcement Learning with Verifiable Rewards (RLVR) activates the model's latent evaluation ability via curriculum learning, producing a high-fidelity endogenous reward signal. Second, Reinforcement Learning with Model-rewarded Thinking (RLMT) leverages this signal to optimize the generative reasoning policy. Experiments show that SEER consistently outperforms state-of-the-art baselines in evaluation accuracy, reprompting efficiency, and generation quality, without sacrificing general multimodal capabilities.

多模态自进化提示工程强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。