arXiv:2602.15460cs.LGcs.CV2026-02被引 1

测试多模态大模型在简单导航任务中的推理泛化能力

On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks

  • 设计网格导航任务,对比文本与视觉输入的推理效果
  • 纯文本模型在跨场景测试中表现优于图像模型
  • 多种文本格式结合的推理路径提升远超预期的泛化能力

将推理机制融入大型语言模型和大型视觉-语言模型,显著提升了其能力。然而,推理模型的泛化能力仍缺乏清晰定义与深入理解。本文提出一个评估框架,严格检验链式思维(CoT)方法在简单规划任务中的泛化性能。具体而言,研究基于网格的导航任务:模型接收地图后需输出一系列移动指令,引导角色从起点到达目标点并避开障碍物。该任务的灵活性支持对不同输入表示(视觉与文本)和CoT推理策略的微调,并在分布内(ID)与分布外(OOD)测试条件下系统评估。实验表明,尽管CoT能提升所有表示形式下的分布内泛化,但在控制与分布内数据的偶然匹配后,大多数情况下的分布外泛化(如更大地图)依然非常有限。令人惊讶的是,结合多种文本格式的推理轨迹展现出最佳且非平凡的分布外泛化能力。最终,纯文本模型始终优于使用图像输入的模型,包括一种依赖潜在空间推理的近期方法。

原文摘要 · Abstract (English)

Integrating reasoning in large language models and large vision-language models has recently led to significant improvement of their capabilities. However, the generalization of reasoning models is still vaguely defined and poorly understood. In this work, we present an evaluation framework to rigorously examine how well chain-of-thought (CoT) approaches generalize on a simple planning task. Specifically, we consider a grid-based navigation task in which a model is provided with a map and must output a sequence of moves that guides a player from a start position to a goal while avoiding obstacles. The versatility of the task and its data allows us to fine-tune model variants using different input representations (visual and textual) and CoT reasoning strategies, and systematically evaluate them under both in-distribution (ID) and out-of-distribution (OOD) test conditions. Our experiments show that, while CoT reasoning improves in-distribution generalization across all representations, out-of-distribution generalization (e.g., to larger maps) remains very limited in most cases when controlling for trivial matches with the ID data. Surprisingly, we find that reasoning traces which combine multiple text formats yield the best (and non-trivial) OOD generalization. Finally, purely text-based models consistently outperform those utilizing image-based inputs, including a recently proposed approach relying on latent space reasoning.

多模态推理泛化链式思维视觉规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。