arXiv:2609.05583cs.CV2026-09综述

让3D模型能用文字或图像检索,支持零样本识别与生成。

An overview of 3D Vision-Language Models

论文配图:An overview of 3D Vision-Language Models
图 1 · 摘自论文原文
  • 用对比学习对齐3D与文本图像嵌入,实现跨模态对齐。
  • 支持零样本分类、跨模态检索和开放词汇识别,性能超越传统方法。
  • 适合研究3D生成、机器人视觉与多模态大模型的学者参考。

视觉语言模型(VLMs)通过对齐视觉与文本嵌入,重塑计算机视觉,使模型能用自然语言理解与推理视觉概念。传统3D深度学习模型通常针对特定任务训练,如分类、分割或检测,难以在嵌入空间中支持以文本或图像为查询的跨模态检索。为此,基于对比语言-图像预训练(CLIP)的方法将3D嵌入与预训练的图像和文本表示对齐,催生了3D视觉语言模型(3D VLMs),支持零样本分类、跨模态检索及3D形状的开放词汇识别。本文综述3D VLMs,涵盖3D表示基础、嵌入编码、跨模态对比对齐、现代多模态框架,以及3D视觉大语言模型(3D VLLMs)。介绍多模态嵌入对齐的对比学习核心定义,并强调语言引导的3D高斯点云渲染、3D形状生成与机器人具身智能等最新进展。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natural language. Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries. To address this issue, Contrastive Language-Image Pretraining (CLIP)-based methods align 3D embeddings with pretrained image and text representations, giving rise to 3D Vision-Language Models (3D VLMs) that support zero-shot classification, cross-modal retrieval, and open-vocabulary recognition of 3D shapes. This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs). We present the main definitions of contrastive learning for multimodal embedding alignment and highlight recent advances in language-guided 3D Gaussian splatting, 3D shape generation, and embodied AI for robotics.

3D视觉多模态大模型生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。