短而精的推理链反而更利于视觉推理泛化
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- 用简洁的接地步骤代替长推理链,提升模型泛化能力
- 实验显示长推理链仅加速收敛,不提高最终性能上限
- 在迷宫任务中,最少接地结果的链路泛化效果最佳
我们研究不同思维链(CoT)设计对视觉语言模型(VLMs)获取可泛化视觉推理能力的影响。尽管长或视觉化的CoT(如‘用图像思考’)被广泛用于监督中间推理过程,但其有效机制尚不明确。为此,我们采用一个受控的迷宫求解基准,其中推理规则完全依赖视觉,难度可通过网格大小调节,且所有中间步骤可自动生成。在Qwen2.5-VL-7B模型上,基于标准SFT-then-RL流程,对比三种代表性CoT格式:语言CoT、接地CoT(含空间坐标轨迹)、视觉CoT(含图像操作)。实验发现:视觉和较长的CoT主要加速收敛,但无法突破最终性能上限;仅包含必要接地步骤的简洁CoT优于长链条;尤为关键的是,仅保留最小接地结果的CoT在不同迷宫尺寸下泛化表现最佳。这些结论在其他视觉中心任务中得到验证。研究揭示了‘短即长’效应,为构建更具泛化能力的SFT数据集提供实用指导。
原文摘要 · Abstract (English)
We study how different Chain-of-Thought (CoT) designs affect the acquisition of the generalizable visual reasoning ability in vision-language models (VLMs). While CoT data, especially long or visual CoT such as "think with image", has been widely used to supervise intermediate reasoning, it remains unclear why specific CoT designs help and which ones truly support generalizable reasoning. To systematically evaluate this, we focus on a controlled maze-solving benchmark where reasoning rules are fully visual, difficulty can be tuned by grid size, and all the intermediate steps can be automatically generated. Using Qwen2.5-VL-7B under a standard SFT-then-RL pipeline, we compare three representative CoT formats: Language CoT, Grounding CoT (with spatial coordinate trajectories), and Visual CoT (with image manipulations). Our experiments reveal that visual and longer CoT mainly accelerate convergence but do not lift the final performance ceiling; concise CoT containing only essential grounding steps outperforms longer traces; and, strikingly, CoT retaining only the minimal grounding results generalizes best across different maze sizes. We further validate these insights on other vision-centric tasks. These findings highlight a "short is long" effect and provide practical guidance for constructing more generalizable SFT datasets for visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。