arXiv:2605.28805cs.CLcs.AI2026-05被引 3

用符号化推理提升多模态模型验证精度,实现精准纠错。

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

论文配图:OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
图 1 · 摘自论文原文
  • 用边界框等符号输出替代文字解释,支持高效规则奖励
  • 分离二元判断与元验证的强化学习目标,性能显著提升
  • 适合需要高可靠性和可解释性的多模态系统部署

视觉输出在多模态大语言模型中日益重要,可靠且细粒度的验证对通用基础模型的扩展至关重要。本文研究多模态元验证,利用验证器生成的推理过程而非仅决策信号,并探索如何有效将元验证反馈融入多模态验证器训练。发现两个关键点:第一,符号化验证输出(如边界框)优于文本解释,可实现基于规则的强化学习奖励,避免依赖辅助裁判模型的模型化奖励;第二,将二元判断与元验证的强化学习目标解耦,显著优于联合优化,因两者输出结构和学习动态存在本质差异。基于此,我们训练了OmniVerifier-M1,一个利用符号化元验证与解耦强化学习的通用视觉验证器。OmniVerifier-M1具备强验证能力与细粒度错误定位能力,并进一步支持M1-TTS,一种由验证器驱动的代理式生成系统,实现区域级动态自纠正。该方法为更可靠、可解释、细粒度的多模态验证开辟路径,助力更安全可控的基础模型部署。

原文摘要 · Abstract (English)

Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In this work, we investigate multimodal meta-verification, which leverages verifier-generated rationales rather than decision-only signals, and explore how to effectively incorporate meta-verification feedback into multimodal verifier training. We identify two key findings. First, symbolic verifier outputs (e.g., bounding boxes) outperform textual explanations as meta-verification rationales, enabling efficient rule-based reinforcement learning rewards while avoiding reliance on model-based rewards from auxiliary judge models. Second, decoupling reinforcement learning objectives for binary judgment and meta-verification substantially outperforms joint reward optimization, due to intrinsic differences in output structure and learning dynamics. Based on these insights, we train OmniVerifier-M1, a generalist visual verifier leveraging symbolic meta-verification and decoupled reinforcement learning. OmniVerifier-M1 provides robust verification and fine-grained error localization, and further enables M1-TTS, a verifier-driven agentic generation system achieving dynamic region-level self-correction. This approach paves the way for more reliable, interpretable, and fine-grained multimodal verification, supporting safer and more controllable foundation model deployment.

多模态验证符号推理强化学习自纠错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。