arXiv:2508.12109cs.CVcs.AI2025-08被引 23

让AI像人一样边看图边思考,提升多模态推理能力

Simple o3: Towards Interleaved Vision-Language Reasoning

  • 通过观察-推理-行动循环生成带可执行视觉操作的推理链
  • 在146K数据集上训练后,在多个基准测试中超越现有方法
  • 首次系统分析不同交互策略,适合研究多模态推理的开发者

多模态大语言模型在视觉语言任务中表现优异,但其在多模态场景下的长链思维(CoT)能力仍待深入探索。受OpenAI o3模型启发,我们提出Simple o3,一个端到端框架,通过监督微调(SFT)将动态工具交互(如裁剪、缩放、复用)融入交错式视觉语言推理。该方法构建了可扩展的数据合成流水线,基于“观察-推理-行动”循环生成高质量交错视觉语言推理链,包含可执行视觉操作与严格验证,形成开源的TWI-Tools-146K数据集。实验表明,Simple o3在多个基准测试中性能显著优于现有方法。通过增强推理能力,Simple o3建立了一种强大且计算成本可控的多模态推理范式。我们首次深入分析不同交错推理策略的影响,发现引入额外视觉标记,并复用和放大原图,显著提升模型视觉推理与细粒度感知;而基于精确视觉定位的图像裁剪,有助于模型聚焦关键实体或区域,进一步增强能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown impressive performance on vision-language tasks, but their long Chain-of-Thought (CoT) capabilities in multimodal scenarios remain underexplored. Inspired by OpenAI's o3 model, which emulates human-like ''thinking with image'' through iterative visual transformations and linguistic reasoning, we propose Simple o3, an end-to-end framework that integrates dynamic tool interactions (e.g., cropping, zooming, and reusing) into interleaved vision-language reasoning via supervised fine-tuning (SFT). Our approach features a scalable data synthesis pipeline that generates high-quality interleaved vision-language reasoning chains via an ''observe-reason-act'' cycle, complete with executable visual operations and rigorous verification, yielding the open-source TWI-Tools-146K dataset. Experimental results demonstrate Simple o3's superior performance on diverse benchmarks, outperforming existing approaches. By combining enhanced reasoning capabilities, Simple o3 establishes a powerful yet computationally affordable paradigm for advancing multimodal reasoning. Remarkably, we provide the first in-depth analysis of different interleaved reasoning strategies, offering insights into their impact on model performance. We found that by introducing additional visual tokens for interleaved vision-language reasoning, reusing and magnifying the original image significantly improves the model's visual reasoning and fine-grained perception, while image cropping based on precise visual grounding allows the model to effectively focus on key entities or regions, further enhancing its capabilities.

多模态推理视觉语言链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。