arXiv:2607.24407cs.CV2026-07中稿 · ACM MM 2026

提出Motto框架,让多模态模型同时精准定位与复杂推理。

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

论文配图:Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
图 1 · 摘自论文原文
  • 用空间对齐的思维标记显式关联视觉位置,提升感知能力。
  • 动态切换接地模式,在推理链中实现跨任务鲁棒定位。
  • 构建新评测基准PR-Bench,专门评估感知与推理的融合效果。

多模态大模型在定位任务上取得显著进展,但现有方法仍难以统一精确定位与复杂推理。文本基方法依赖坐标或索引预测,严重限制模型对密集视觉对象的感知能力;而基于潜在标记的方法使用无空间参考的特殊标记,解码机制缺乏思考步骤,削弱了高层推理能力。为此,我们提出混合思维标记(Mixture-of-Thought-Tokens, Motto),一种新的自由形式多模态定位方法,弥合感知与推理之间的鸿沟,使多模态大模型能应对多样且任意的定位查询。具体而言,我们引入空间对齐的思维标记化,显式将特殊标记与空间位置对齐,确保清晰的空间对应关系和视觉可解释性。同时设计上下文自适应的思维链标记机制,在交错推理链中动态切换定位模式,实现不同复杂度任务下的鲁棒定位。此外,我们构建了PR-Bench——一个新的指代表达理解基准,用于评估感知与推理的差距。大量实验表明,Motto在多种自由形式定位任务中达到领先性能。

原文摘要 · Abstract (English)

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model for dense visual objects. Meanwhile, latent token-based methods employ special tokens without inherent spatial references and use a decoding mechanism that lacks thinking steps, weakening high-level reasoning capabilities. Consequently, developing a unified framework that excels in both perception and reasoning remains challenging. To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap, enabling MLLMs to empower diverse, arbitrary grounding queries. Specifically, we introduce Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability. We further design a Context-Adaptive Chain-of-Tokens that dynamically switch grounding modes within an interleaved reasoning chain, achieving robust grounding across tasks of varying complexity. In addition, we construct PR-Bench, a new referring expression comprehension benchmark to evaluate the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.

多模态定位推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。