arXiv:2503.06232cs.CLcs.CV2025-03被引 6

用思维链提升3D视觉语言对齐,让模型更懂物体形状功能

Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning

  • 在3D视觉语言任务中引入结构化思维链推理
  • 思维链使3D语义定位准确率显著提升,大推理模型效果更优
  • 适合研究多模态推理与具身智能的学者参考

思维链(CoT)推理在自然语言任务中表现优异,但在多模态对齐领域仍待探索。本研究将结构化思维链融入3D视觉语言学习,通过嵌入推理过程增强对齐训练。我们构建了3D-CoT基准数据集,包含层次化思维链标注,覆盖形状识别、功能推断和因果推理。在大型推理模型(LRMs)与大型语言模型(LLMs)上进行对比实验,采用双层评估框架衡量中间推理与最终推理质量。结果表明,使用思维链可显著提升3D语义定位性能,且LRMs比LLMs更有效利用思维链。此外,标注结构影响性能:显式推理标记有助于LLMs,而未标记的思维链更契合LRM的推理模式。分析显示,思维链对多模态推理至关重要,其影响可推广至非3D任务。数据集将公开于https://huggingface.co/datasets/Battam/3D-CoT。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning has proven effective in natural language tasks but remains underexplored in multimodal alignment. This study investigates its integration into 3D vision-language learning by embedding structured reasoning into alignment training. We introduce the 3D-CoT Benchmark, a dataset with hierarchical CoT annotations covering shape recognition, functional inference, and causal reasoning. Through controlled experiments, we compare CoT-structured and standard textual annotations across large reasoning models (LRMs) and large language models (LLMs). Our evaluation employs a dual-layer framework assessing both intermediate reasoning and final inference quality. Extensive experiments demonstrate that CoT significantly improves 3D semantic grounding, with LRMs leveraging CoT more effectively than LLMs. Furthermore, we highlight that annotation structure influences performance-explicit reasoning markers aid LLMs, while unmarked CoT better aligns with LRM inference patterns. Our analyses suggest that CoT is crucial for enhancing multimodal reasoning, with implications beyond 3D tasks. The dataset will be publicly available at https://huggingface.co/datasets/Battam/3D-CoT

3D视觉思维链多模态对齐推理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。