arXiv:2506.08691cs.CV2025-06ACL被引 9

用树搜索和自奖励机制提升视觉语言模型的推理能力。

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism

  • 构建推理树,逐步搜索最优解法路径。
  • 在三个数学推理基准上达到当前最佳表现。
  • 无需额外模型,适合想提升推理效果的研究者。

大型视觉语言模型(LVLMs)在多模态任务中表现出色,但在复杂视觉推理方面仍受限,尤其在使用思维链提示时。本文提出VReST,一种无需训练的新方法,通过蒙特卡洛树搜索与自奖励机制增强推理能力。VReST构建推理树,每个节点代表一个推理步骤,每条路径对应完整推理序列。其创新的多模态自奖励机制结合子问题效用、答案正确性及视觉-语言线索相关性,自动评估推理质量,无需额外模型。VReST超越现有提示方法,在三个多模态数学推理基准上取得当前最优性能。同时验证了测试时缩放定律在多模态任务中的有效性,为未来研究提供新方向。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown exceptional performance in multimodal tasks, but their effectiveness in complex visual reasoning is still constrained, especially when employing Chain-of-Thought prompting techniques. In this paper, we propose VReST, a novel training-free approach that enhances Reasoning in LVLMs through Monte Carlo Tree Search and Self-Reward mechanisms. VReST meticulously traverses the reasoning landscape by establishing a search tree, where each node encapsulates a reasoning step, and each path delineates a comprehensive reasoning sequence. Our innovative multimodal Self-Reward mechanism assesses the quality of reasoning steps by integrating the utility of sub-questions, answer correctness, and the relevance of vision-language clues, all without the need for additional models. VReST surpasses current prompting methods and secures state-of-the-art performance across three multimodal mathematical reasoning benchmarks. Furthermore, it substantiates the efficacy of test-time scaling laws in multimodal tasks, offering a promising direction for future research.

视觉推理树搜索自奖励多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。