arXiv:2412.14835cs.CLcs.AI2024-12被引 34

用主动检索+树搜索提升多模态模型的逐步推理能力

Progressive Multimodal Reasoning via Active Retrieval

  • 通过主动检索从多模态语料中获取关键线索
  • 在三个基准上显著提升多模态模型推理准确率
  • 适合需要复杂逻辑推理的视觉语言任务研究者

多步多模态推理任务对多模态大语言模型(MLLMs)构成重大挑战,现有方法难以有效提升其表现。本文提出AR-MCTS框架,结合主动检索(AR)与蒙特卡洛树搜索(MCTS),逐步增强MLLMs的推理能力。首先构建统一检索模块,从混合模态语料库中提取解决复杂问题的关键支持信息。为弥补自动化多模态推理验证的不足,采用融合主动检索机制的MCTS算法,实现分步标注的自动生成。该策略在每一步动态检索关键见解,突破传统束搜索采样局限,提升推理空间的多样性与可靠性。此外,引入过程奖励模型,逐步对齐以支持多模态推理任务的自动验证。在三个复杂多模态推理基准上的实验结果表明,AR-MCTS可有效提升多种多模态模型性能。进一步分析显示,该方法能优化采样多样性与准确性,实现可靠多模态推理。

原文摘要 · Abstract (English)

Multi-step multimodal reasoning tasks pose significant challenges for multimodal large language models (MLLMs), and finding effective ways to enhance their performance in such scenarios remains an unresolved issue. In this paper, we propose AR-MCTS, a universal framework designed to progressively improve the reasoning capabilities of MLLMs through Active Retrieval (AR) and Monte Carlo Tree Search (MCTS). Our approach begins with the development of a unified retrieval module that retrieves key supporting insights for solving complex reasoning problems from a hybrid-modal retrieval corpus. To bridge the gap in automated multimodal reasoning verification, we employ the MCTS algorithm combined with an active retrieval mechanism, which enables the automatic generation of step-wise annotations. This strategy dynamically retrieves key insights for each reasoning step, moving beyond traditional beam search sampling to improve the diversity and reliability of the reasoning space. Additionally, we introduce a process reward model that aligns progressively to support the automatic verification of multimodal reasoning tasks. Experimental results across three complex multimodal reasoning benchmarks confirm the effectiveness of the AR-MCTS framework in enhancing the performance of various multimodal models. Further analysis demonstrates that AR-MCTS can optimize sampling diversity and accuracy, yielding reliable multimodal reasoning.

多模态推理主动检索树搜索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。