通过比值对比学习提升3D点云的多模态理解精度
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
- 用高密度点云和困难负样本增强视觉文本对齐
- 引入比值作为辅助项,在对比学习中提升区分度
- 适合做3D场景理解与跨模态推理的研究者
近期研究显示,大语言模型不仅限于文本任务,还能处理音频、图像、视频等多模态数据。特别是3D大模型(3D LMMs)因可处理点云等高维数据而取得显著进展。然而,现有训练数据中的视觉与文本内容信息粒度低、表述模糊,制约了跨模态理解的精度。为此,我们提出CL3DOR:基于高分辨率点云的比值对比学习方法,旨在提升视觉与文本内容的特异性与清晰度。具体地,通过增加每物体点云密度,并构建具有信息量的困难负样本以惩罚错误响应。为有效利用这些负样本,我们在传统语言建模损失中引入比值作为辅助项,实现更精准的对比学习。CL3DOR在3D场景理解与推理基准测试中达到当前最优性能。大量实验验证了其核心组件的有效性。
原文摘要 · Abstract (English)
Recent research has demonstrated that Large Language Models (LLMs) are not limited to text-only tasks but can also function as multimodal models across various modalities, including audio, images, and videos. In particular, research on 3D Large Multimodal Models (3D LMMs) is making notable strides, driven by the potential of processing higher-dimensional data like point clouds. However, upon closer examination, we find that the visual and textual content within each sample of existing training datasets lacks both high informational granularity and clarity, which serve as a bottleneck for precise cross-modal understanding. To address these issues, we propose CL3DOR, Contrastive Learning for 3D large multimodal models via Odds ratio on high-Resolution point clouds, designed to ensure greater specificity and clarity in both visual and textual content. Specifically, we increase the density of point clouds per object and construct informative hard negative responses in the training dataset to penalize unwanted responses. To leverage hard negative responses, we incorporate the odds ratio as an auxiliary term for contrastive learning into the conventional language modeling loss. CL3DOR achieves state-of-the-art performance in 3D scene understanding and reasoning benchmarks. Additionally, we demonstrate the effectiveness of CL3DOR's key components through extensive experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。