arXiv:2502.13383cs.CLcs.CV2025-02ACL被引 50

用思维链验证提升多模态推理能力,效果超越GPT-4o。

MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification

  • 通过模拟树搜索与拒绝采样生成高质量多模态思维链数据。
  • 在MathVista上达到65.3%准确率,超过GPT-4o的63.8%。
  • 适用于需要高精度多模态推理的研究与应用者。

根据测试时扩展原则,将外部慢思考与验证机制结合已被证明可增强大语言模型的多轮推理能力。然而,在多模态领域,仍缺乏强有力的多模态验证器。本文提出MM-Verifier与MM-Reasoner,通过更长的推理过程和更鲁棒的验证来提升多模态推理性能。首先,我们设计了一种两步式多模态验证数据合成方法,结合基于模拟的树搜索与验证,并采用拒绝采样生成高质量思维链(Chain-of-Thought, COT)数据,用于微调验证模型MM-Verifier。此外,我们提出一种更高效的多模态思维链数据合成方法,弥合文本与多模态推理间的差距。该数据用于微调MM-Reasoner。MM-Verifier在MathCheck、MathVista和MathVerse基准上优于所有更大模型。同时,MM-Reasoner表现出强有效性与可扩展性,性能随数据量增加而提升。最终,结合MM-Reasoner与MM-Verifier的方法在MathVista上取得65.3%准确率,高于使用12次回滚的GPT-4o(63.8%)。

原文摘要 · Abstract (English)

According to the Test-Time Scaling, the integration of External Slow-Thinking with the Verify mechanism has been demonstrated to enhance multi-round reasoning in large language models (LLMs). However, in the multimodal (MM) domain, there is still a lack of a strong MM-Verifier. In this paper, we introduce MM-Verifier and MM-Reasoner to enhance multimodal reasoning through longer inference and more robust verification. First, we propose a two-step MM verification data synthesis method, which combines a simulation-based tree search with verification and uses rejection sampling to generate high-quality Chain-of-Thought (COT) data. This data is then used to fine-tune the verification model, MM-Verifier. Additionally, we present a more efficient method for synthesizing MMCOT data, bridging the gap between text-based and multimodal reasoning. The synthesized data is used to fine-tune MM-Reasoner. Our MM-Verifier outperforms all larger models on the MathCheck, MathVista, and MathVerse benchmarks. Moreover, MM-Reasoner demonstrates strong effectiveness and scalability, with performance improving as data size increases. Finally, our approach achieves strong performance when combining MM-Reasoner and MM-Verifier, reaching an accuracy of 65.3 on MathVista, surpassing GPT-4o (63.8) with 12 rollouts.

多模态推理思维链验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。