让文字和图像交替推理,实现更智能的多模态思考。
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- 文字与图像交替推理,互为补充而非重复
- 在视觉任务上提升34.7%,超越更大模型
- 能自适应切换推理模式,具备新视觉操作能力
多模态推理需要语言与视觉的迭代协同,但有意义的交错思维链仍不明确。我们提出文本与图像思维应互补而非同构,据此构建了ThinkMorph——一个在约2.4万条高质量交错推理轨迹上微调的统一模型,涵盖不同视觉参与度的任务。ThinkMorph学会生成逐步推进的图文推理步骤,在具体操作视觉内容的同时保持连贯的语义逻辑。其在以视觉为中心的基准上平均超越基线模型34.7%,并能泛化至域外任务,表现媲美甚至超过更大的专有视觉语言模型(VLM)。此外,ThinkMorph展现出涌现的多模态智能,包括未见的视觉操作能力、推理模式的自适应切换,以及通过多样化多模态思维实现更优的测试时扩展。这些结果为统一模型在多模态推理中的涌现能力提供了新方向。
原文摘要 · Abstract (English)
Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what constitutes a meaningful interleaved chain of thought. We posit that text and image thoughts should function as complementary rather than isomorphic modalities that mutually advance reasoning. Guided by this principle, we build ThinkMorph, a unified model fine-tuned on approximately 24K high-quality interleaved reasoning traces spanning tasks with varying visual engagement. ThinkMorph learns to generate progressive text-image reasoning steps that concretely manipulate visual content while maintaining coherent verbal logic. It delivers large gains on vision-centric benchmarks (averaging 34.7 percent over the base model) and generalizes to out-of-domain tasks, matching or surpassing larger and proprietary VLMs. Beyond performance, ThinkMorph exhibits emergent multimodal intelligence, including unseen visual manipulation skills, adaptive switching between reasoning modes, and better test-time scaling through diversified multimodal thoughts. These findings suggest promising directions for characterizing the emergent capabilities of unified models for multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。