arXiv:2510.13804cs.CVcs.AI2025-10被引 17

让AI自检视觉输出,提升多模态推理可靠性

Generative Universal Verifier as Multimodal Meta-Reasoner

  • 构建通用视觉验证器,实现生成过程中的自我反思与优化
  • 在ViVerBench上提升8.3分,验证器具备可泛化的三类原子能力
  • 适用于图像生成、编辑及复杂推理,适合追求高可信度AI系统的研究者

我们提出生成式通用验证器(Generative Universal Verifier),一种面向下一代视觉语言模型与统一多模态模型的新概念与插件,赋予模型在推理与生成过程中对视觉结果进行反思与优化的核心能力。本文主要贡献有三:(1) 构建覆盖16类关键任务的ViVerBench基准,结果表明现有视觉语言模型在这些任务中持续表现不佳,与人类水平存在显著差距;(2) 设计两条自动化数据构建管道,训练出首个全能型生成式验证器OmniVerifier-7B,其在ViVerBench上取得+8.3的显著提升,并识别出视觉验证中的三种基础能力及其协同机制;(3) 提出OmniVerifier-TTS,一种序列化测试时扩展范式,利用通用验证器在统一模型中桥接图像生成与编辑,通过迭代微调提升生成上限。实验显示,其在T2I-ReasonBench和GenEval++上分别提升+3.7和+4.3,优于Best-of-N等并行方法。该工作为多模态推理提供了可靠视觉验证能力,推动生成过程中的可信反思与可扩展测试时优化,迈向更可信、可控的下一代推理系统。

原文摘要 · Abstract (English)

We introduce Generative Universal Verifier, a novel concept and plugin designed for next-generation multimodal reasoning in vision-language models and unified multimodal models, providing the fundamental capability of reflection and refinement on visual outcomes during the reasoning and generation process. This work makes three main contributions: (1) We build ViVerBench, a comprehensive benchmark spanning 16 categories of critical tasks for evaluating visual outcomes in multimodal reasoning. Results show that existing VLMs consistently underperform across these tasks, underscoring a substantial gap from human-level capability in reliable visual verification. (2) We design two automated pipelines to construct large-scale visual verification data and train OmniVerifier-7B, the first omni-capable generative verifier trained for universal visual verification and achieves notable gains on ViVerBench(+8.3). Through training, we identify three atomic capabilities in visual verification and demonstrate how they generalize and interact synergistically. (3) We propose OmniVerifier-TTS, a sequential test-time scaling paradigm that leverages the universal verifier to bridge image generation and editing within unified models, enhancing the upper bound of generative ability through iterative fine-grained optimization. Beyond generation, we extend universal verifier to broader world-modeling interleaved reasoning scenarios. Empirically, OmniVerifier-TTS achieves improvements on T2I-ReasonBench(+3.7), and GenEval++(+4.3), outperforming existing parallel test-time scaling methods, such as Best-of-N. By endowing multimodal reasoning with reliable visual verification, OmniVerifier advances both reliable reflection during generation and scalable test-time refinement, marking a step toward more trustworthy and controllable next-generation reasoning systems.

多模态推理视觉验证生成优化测试时扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。