无需训练,用动态视觉草图提升多模态模型空间推理能力
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
- 将动态视觉草图叠加在输入图像上,辅助文本思维链推理
- 在新构建的GRASSLAND基准上,性能显著优于传统方法
- 无需微调模型,适合快速部署于各类动态空间任务
尽管思维链(CoT)已推动多模态大模型在复杂推理中的发展,现有方法仍局限于文本或静态视觉领域,在动态空间推理任务中表现不佳。为此,我们提出了GRASSLAND——一个全新的迷宫导航基准,用于评估动态空间推理能力。实验表明,将动态视觉草图与文本推理链结合,能显著超越传统方法,为动态环境下的空间推理提供新视角。为推广此能力,我们提出D2R(Dynamic Draft-Augmented Reasoning),一种无需训练的框架,可无缝将文本CoT与对应视觉草图集成至多模态大模型。大量实验显示,D2R在多种任务中持续提升性能,无需模型微调即可建立动态空间推理的可靠基线。项目开源地址:https://github.com/Cratileo/D2R。
原文摘要 · Abstract (English)
While chains-of-thought (CoT) have advanced complex reasoning in multimodal large language models (MLLMs), existing methods remain confined to text or static visual domains, often faltering in dynamic spatial reasoning tasks. To bridge this gap, we present GRASSLAND, a novel maze navigation benchmark designed to evaluate dynamic spatial reasoning. Our experiments show that augmenting textual reasoning chains with dynamic visual drafts, overlaid on input images, significantly outperforms conventional approaches, offering new insights into spatial reasoning in evolving environments. To generalize this capability, we propose D2R (Dynamic Draft-Augmented Reasoning), a training-free framework that seamlessly integrates textual CoT with corresponding visual drafts into MLLMs. Extensive evaluations demonstrate that D2R consistently enhances performance across diverse tasks, establishing a robust baseline for dynamic spatial reasoning without requiring model fine-tuning. Project is open at https://github.com/Cratileo/D2R.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。