arXiv:2605.22013cs.CVcs.GR2026-05International Conf…被引 2

让3D点云理解具备逻辑推理能力,提升模型泛化性。

PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought

论文配图:PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought
图 1 · 摘自论文原文
  • 构建基于思维链的3D点云指令数据集,增强模型推理能力。
  • 在5.5万条数据上训练,显著提升分类与描述生成性能。
  • 适合需要精准3D理解的工业建模与机器人导航场景。

通过语言理解3D点云仍是计算机图形学与视觉计算中的核心挑战,主要源于点云数据结构不规则及现有多模态模型缺乏显式推理能力。尽管思维链(Chain-of-Thought, CoT)在大语言模型和图像多模态模型中表现优异,但其在3D理解中的应用仍基本空白。本文提出一种以数据为中心的框架,构建大规模面向3D点云理解的思维链监督数据。该框架采用两阶段流程:第一阶段利用视觉-语言模型评估并优化点文本指令数据质量,结合参考信息进行精细化修正;第二阶段通过人机协同提示优化(HiLPO)生成高质量推理路径。基于此方法,我们构建了包含55,000个样本的PoCoTI数据集,其中包含明确的推理路径。在此基础上微调PointLLM得到PointLLM-R,一个具备推理能力的3D多模态语言模型。大量实验表明,PointLLM-R在生成式3D分类与描述任务中达到当前最优性能,并在真实扫描点云和多轮对话场景中表现出强泛化能力。

原文摘要 · Abstract (English)

Understanding 3D point clouds through language remains a fundamental challenge in computer graphics and visual computing, due to the irregular structure of point cloud data and the lack of explicit reasoning in existing 3D multimodal models. While Chain-of-Thought (CoT) reasoning has shown strong effectiveness in LLMs and image-based MLLMs, its extension to 3D understanding remains largely underexplored. In this paper, we propose a data-centric framework for constructing large-scale CoT supervision tailored to 3D point cloud understanding. Our framework consists of a two-stage pipeline that first refines point-text instruction data via vision-language-model-based quality evaluation and reference-guided refinement, and then synthesizes high-quality reasoning paths through Human-in-the-Loop Prompt Optimization (HiLPO). Using this approach, we build PoCoTI, a CoT-enhanced point-text instruction-following dataset containing 55K samples with explicit reasoning paths. Fine-tuning PointLLM on PoCoTI yields PointLLM-R, a reasoning-capable 3D multimodal language model. Extensive experiments on generative 3D classification and captioning demonstrate that PointLLM-R achieves state-of-the-art performance and generalizes robustly to real-world scanned point clouds and multi-turn dialogue scenarios.

3D理解思维链点云多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。