通过细粒度偏好优化,提升视觉语言模型的空间推理能力。
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
- 用多模型蒙特卡洛树搜索生成多样逻辑链推理路径。
- 细粒度偏好优化使空间定性与定量任务分别提升9.0%和4.1%。
- 适合需要精准空间理解的视觉推理场景,如机器人导航。
当前视觉语言模型在需要多步逻辑与精确空间对齐的细粒度空间推理任务中表现不佳。本文提出SpatialReasoner-R1模型,通过设计多模型蒙特卡洛树搜索(M3CTS)方法,生成多样且逻辑一致的长链思维(LongCOT)推理轨迹以构建高质量监督信号。进一步提出细粒度直接偏好优化(fDPO),引入分段偏好粒度,结合空间奖励机制(评估视觉一致性、空间定位与逻辑连贯性),指导描述性定位与逻辑推理。实验表明,fDPO在空间定性与定量任务上相对标准DPO分别提升4.1%与9.0%。基于fDPO训练的SpatialReasoner-R1在SpatialRGPT-Bench上达到新SOTA,平均准确率领先最强基线9.4%,同时保持通用视觉语言任务竞争力。
原文摘要 · Abstract (English)
Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high-quality supervision for spatial reasoning, we design a Multi-Model Monte Carlo Tree Search (M3CTS) method that generates diverse, logically consistent Long Chain-of-Thought (LongCOT) reasoning trajectories. In addition, we propose a fine-grained Direct Preference Optimization (fDPO) method that introduces segment-specific preference granularity for descriptive grounding and logical reasoning, guided by a spatial reward mechanism that evaluates candidate responses based on visual consistency, spatial grounding, and logical coherence. Experimental results demonstrate that fDPO achieves relative performance gains of 4.1% and 9.0% over standard DPO on spatial qualitative and quantitative tasks, respectively. SpatialReasoner-R1, trained with fDPO, sets a new SoTA on SpatialRGPT-Bench, outperforming the strongest baseline by 9.4% in average accuracy, while maintaining competitive performance on general vision-language tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。