arXiv:2512.10941cs.CVcs.AI2025-12中稿 · CVPR被引 13

让模型用统一的隐变量自由跨模态思考,提升空间推理能力。

Mull-Tokens: Modality-Agnostic Latent Thinking

  • 设计通用隐变量令牌,可承载图像与文本的中间思考信息。
  • 在四类空间推理任务中平均提升3%,最高达16%。
  • 适合需要多模态抽象推理的研究者与工程师。

推理不仅限于语言;现实世界需要对空间、时间、可操作性等进行推理,仅靠文字无法表达。现有多模态模型在图像推理方面脆弱且难以扩展,依赖专用工具调用、昂贵的图像生成或手工制作的推理数据来切换文本与图像思维。为此,我们提出更简单的方案——Mull-Tokens,一种模态无关的隐变量令牌,通过预训练掌握图像或文本模态中的中间信息,使模型能自由地向正确答案推进。我们基于隐式推理框架探索最佳训练实践:首先利用交错文本-图像轨迹进行监督训练,然后仅用最终答案进行无监督微调。在涉及解谜和视角转换的四个挑战性空间推理基准上,实验表明,Mull-Tokens优于多个仅使用文本推理或交错图文推理的基线模型,在平均表现上提升3%,在以解谜为主的复杂子集上最高提升16%。该工作为文本与视觉推理的对齐难题提供了一种简洁有效的解决方案。

原文摘要 · Abstract (English)

Reasoning goes beyond language; the real world requires reasoning about space, time, affordances, and much more that words alone cannot convey. Existing multimodal models exploring the potential of reasoning with images are brittle and do not scale. They rely on calling specialist tools, costly generation of images, or handcrafted reasoning data to switch between text and image thoughts. Instead, we offer a simpler alternative -- Mull-Tokens -- modality-agnostic latent tokens pre-trained to hold intermediate information in either image or text modalities to let the model think free-form towards the correct answer. We investigate best practices to train Mull-Tokens inspired by latent reasoning frameworks. We first train Mull-Tokens using supervision from interleaved text-image traces, and then fine-tune without any supervision by only using the final answers. Across four challenging spatial reasoning benchmarks involving tasks such as solving puzzles and taking different perspectives, we demonstrate that Mull-Tokens improve upon several baselines utilizing text-only reasoning or interleaved image-text reasoning, achieving a +3% average improvement and up to +16% on a puzzle solving reasoning-heavy split compared to our strongest baseline. Adding to conversations around challenges in grounding textual and visual reasoning, Mull-Tokens offers a simple solution to abstractly think in multiple modalities.

多模态推理隐变量空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。