arXiv:2601.08811cs.CVcs.AI2026-01

用自动生成的带推理过程数据训练出更强的3D视觉定位模型

Reasoning Matters for 3D Visual Grounding

  • 自动合成含推理链的3D视觉定位数据
  • 仅用1.6%数据超越此前顶尖模型
  • 适合关注3D理解与大模型推理的开发者

大型语言模型(LLM)强大的推理能力推动了数学、编程和科学发现等领域的研究。然而,3D视觉定位作为3D理解的基础任务,仍因现有模型推理能力有限而面临挑战。当前多数方法依赖文本编码器与视觉特征编码器融合跨模态特征进行目标预测,且需大量3D标注数据监督训练。尽管有研究尝试通过扩大合成数据来训练更强的3D视觉定位大模型,但性能提升有限且不随数据成本线性增长。本文提出一种3D视觉定位数据生成管道,可自动合成带相应推理过程的3D视觉定位数据。我们利用生成数据对LLM进行微调,提出Reason3DVG-8B模型,在仅使用3D-GRAND 1.6%训练数据的情况下,性能超越其基线,验证了所提数据的有效性及推理在3D视觉定位中的关键作用。

原文摘要 · Abstract (English)

The recent development of Large Language Models (LLMs) with strong reasoning ability has driven research in various domains such as mathematics, coding, and scientific discovery. Meanwhile, 3D visual grounding, as a fundamental task in 3D understanding, still remains challenging due to the limited reasoning ability of recent 3D visual grounding models. Most of the current methods incorporate a text encoder and visual feature encoder to generate cross-modal fuse features and predict the referring object. These models often require supervised training on extensive 3D annotation data. On the other hand, recent research also focus on scaling synthetic data to train stronger 3D visual grounding LLM, however, the performance gain remains limited and non-proportional to the data collection cost. In this work, we propose a 3D visual grounding data pipeline, which is capable of automatically synthesizing 3D visual grounding data along with corresponding reasoning process. Additionally, we leverage the generated data for LLM fine-tuning and introduce Reason3DVG-8B, a strong 3D visual grounding LLM that outperforms previous LLM-based method 3D-GRAND using only 1.6% of their training data, demonstrating the effectiveness of our data and the importance of reasoning in 3D visual grounding.

3D视觉定位大模型推理合成数据多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。