arXiv:2511.00054cs.LGcs.AI2025-11中稿 · NeurIPS

用自动验证生成高质量空间推理数据,助力小模型高效学习。

SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation

  • 通过自动验证器生成多步、多工具的高保真推理链。
  • 在CLEVR-Humans上提升推理质量17%,方差降低超40%。
  • 适合想让小模型掌握复杂空间推理的研究者使用。

尽管视觉语言模型在诸多任务中表现优异,但在需要问题分解与策略性工具使用的复杂空间推理上仍显不足。微调小型可部署模型是实现高性能的有效路径,但受限于高质量、分步推理数据的缺乏。为此,我们提出SpatialTraceGen框架,将大模型的推理过程蒸馏为高质量的多跳、多工具推理轨迹数据集。其核心创新是自动化验证器,可规模化保证每一步推理的准确性,替代昂贵的人工标注。在CLEVR-Humans基准测试中,该验证器引导流程使轨迹平均质量评分提升17%,质量方差下降超过40%。SpatialTraceGen生成专家级推理轨迹数据,为小模型微调和样本高效的离线强化学习提供结构化、分步示例。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) excel in many areas, they struggle with complex spatial reasoning, which requires problem decomposition and strategic tool use. Fine-tuning smaller, more deployable models offers an efficient path to strong performance, but this is hampered by a major bottleneck: the absence of high-quality, step-by-step reasoning data. To address this data-efficiency gap, we introduce SpatialTraceGen, a framework to distill the reasoning processes of a large teacher model into a high-quality dataset of multi-hop, multi-tool reasoning traces. A key innovation is our automated Verifier, which scalably ensures the fidelity of each reasoning step, providing a cost-effective alternative to manual human annotation. On the CLEVR-Humans benchmark, this verifier-guided process improves the average quality score of traces by 17\% while reducing quality variance by over 40\%. SpatialTraceGen delivers a dataset of expert traces, providing the structured, step-by-step examples of tool use necessary for effective fine-tuning and sample-efficient offline reinforcement learning.

空间推理知识蒸馏自动验证数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。