让大模型根据问题复杂度自动在语言和网格间切换,提升空间推理准确率。
Spatial Reasoning via Modality Switching Between Language and Symbolic Representation

- 基于可信度与复杂度信号设计模态切换机制
- 在网格表示下模型性能最高提升42%
- 适合需要空间推理的逻辑题与规划任务
人类推理天然多模态:当问题变难时,我们很少仅靠语言思考,常通过画图或网格来外化思维、理清结构、避免错误。本文研究两个问题:(a) 将多跳文本-空间故事转化为布局或网格等几何感知模态,是否比纯语言推理更优;(b) 模型能否自主决定何时用语言推理,何时切换到结构化模态。为此,我们提出一种基于可信度与复杂度信号的切换指标,预测何时将空间故事转为结构表示可提升性能。实验表明,在各类设置中,从语言推理切换至网格表示,模型表现最高提升42%,首次实现大模型推理中的有原则的模态选择。该工作揭示了模态选择对推理结果的关键影响。
原文摘要 · Abstract (English)
Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared to natural language-based inference; and (b) whether a model can decide when to rely on natural language reasoning and when to switch to a structured modality. We address these questions by introducing a switching metric based on trustworthiness and complexity signals, which estimates when grounding a spatial story into structure is likely to improve performance. This takes a first step toward principled modality selection in Large Language Model (LLM) reasoning. Across our settings, switching from natural language-based reasoning to a grid-based representation improves LLM performance by up to 42%, highlighting the importance of modality choice in shaping reasoning outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。