让视觉语言模型像人一样一步步看图思考,提升复杂推理能力。
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
- 用多模态树搜索实现边看图边推理的渐进式思考过程。
- 无需微调,在几何与空间推理任务上达到当前最优性能。
- 适合需要深度视觉推理的AI研究者与开发者参考。
近年来,大型视觉语言模型展现了卓越的能力,但在涉及复杂推理的任务中常表现不佳,而这类任务人类通常通过视觉辅助和逐步思考来解决。现有方法虽尝试文本慢思考或简单视觉协助,却未能捕捉人类视觉-语言推理中交错复杂的本质。受人类慢思考机制启发,本文提出VisuoThink框架,无缝融合视觉空间与语言领域,通过测试时的前瞻树搜索实现多模态慢思考,支持渐进式视觉-文本推理。大量实验表明,该方法在不进行微调的情况下,仅通过推理阶段扩展即显著提升模型推理能力,在几何与空间推理任务上达到当前最优表现。
原文摘要 · Abstract (English)
Recent advancements in Large Vision-Language Models have showcased remarkable capabilities. However, they often falter when confronted with complex reasoning tasks that humans typically address through visual aids and deliberate, step-by-step thinking. While existing methods have explored text-based slow thinking or rudimentary visual assistance, they fall short of capturing the intricate, interleaved nature of human visual-verbal reasoning processes. To overcome these limitations and inspired by the mechanisms of slow thinking in human cognition, we introduce VisuoThink, a novel framework that seamlessly integrates visuospatial and linguistic domains. VisuoThink facilitates multimodal slow thinking by enabling progressive visual-textual reasoning and incorporates test-time scaling through look-ahead tree search. Extensive experiments demonstrate that VisuoThink significantly enhances reasoning capabilities via inference-time scaling, even without fine-tuning, achieving state-of-the-art performance in tasks involving geometry and spatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。