arXiv:2511.12079cs.CV2025-11AAAI

用文本提示引导点云量化,提升语义理解能力

Point Cloud Quantization through Multimodal Prompting for 3D Understanding

  • 用预训练模型的文本嵌入做原型先验,增强代表性
  • 多模态提示自适应优化原型,减少视觉语言语义鸿沟
  • 在ModelNet40和ScanObjectNN上表现更优,适合3D理解任务

向量量化已成为大规模多模态模型中统一异构表示的强大工具,通过离散标记编码实现。然而其效果依赖于稳健的码本设计。当前基于原型的方法依赖可训练向量或聚类中心,在代表性和可解释性方面仍有不足,尽管多模态对齐在视觉-语言模型中展现出潜力。为解决这些局限,我们提出一种简单而有效的多模态提示驱动的点云量化框架。该方法基于两个核心洞察:1)预训练模型的文本嵌入通过多对一对比对齐天然蕴含视觉语义,可自然作为鲁棒的原型先验;2)多模态提示可自适应优化这些原型,有效缓解视觉-语言语义鸿沟。框架引入双约束量化空间,通过紧凑性与分离性正则化,无缝融合视觉与原型特征,生成联合编码几何与语义信息的混合表示。此外,采用Gumbel-Softmax松弛实现可微离散化,同时保持量化稀疏性。在ModelNet40和ScanObjectNN数据集上的大量实验明确证明了所提方法的优越性。

原文摘要 · Abstract (English)

Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current prototype-based approaches relying on trainable vectors or clustered centroids fall short in representativeness and interpretability, even as multimodal alignment demonstrates its promise in vision-language models. To address these limitations, we propose a simple multimodal prompting-driven quantization framework for point cloud analysis. Our methodology is built upon two core insights: 1) Text embeddings from pre-trained models inherently encode visual semantics through many-to-one contrastive alignment, naturally serving as robust prototype priors; and 2) Multimodal prompts enable adaptive refinement of these prototypes, effectively mitigating vision-language semantic gaps. The framework introduces a dual-constrained quantization space, enforced by compactness and separation regularization, which seamlessly integrates visual and prototype features, resulting in hybrid representations that jointly encode geometric and semantic information. Furthermore, we employ Gumbel-Softmax relaxation to achieve differentiable discretization while maintaining quantization sparsity. Extensive experiments on the ModelNet40 and ScanObjectNN datasets clearly demonstrate the superior effectiveness of the proposed method.

点云量化多模态提示3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。