用动态体素与元嵌入提升3D点云的语言感知能力
NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
- 动态多尺度体素根据几何复杂度自适应调整粒度
- 体素特征提取效率提升,重建保真度保持高位
- 适合需要高效3D语义理解的机器人与自动驾驶场景
视觉语言模型和多模态大语言模型的进展推动了以语言为导向的3D场景理解。然而,现有3D语言模型在处理稀疏、大规模点云时面临特征提取慢、表示精度不足的问题。为此,我们提出NeuroVoxel-LM框架,融合神经辐射场(NeRF)与动态分辨率体素化及轻量级元嵌入。具体地,提出动态分辨率多尺度体素化(DR-MSV)技术,依据几何与结构复杂度自适应调整体素粒度,降低计算开销的同时保持重建保真度。此外,提出基于注意力加权与残差融合的词元级自适应池化(TAP-LME)机制,增强从NeRF权重中捕捉细粒度语义的能力。实验表明,DR-MSV显著提升点云特征提取效率与准确率,TAP-LME优于传统最大池化,在细粒度语义建模上表现更优。
原文摘要 · Abstract (English)
Recent breakthroughs in Visual Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have significantly advanced 3D scene perception towards language-driven cognition. However, existing 3D language models struggle with sparse, large-scale point clouds due to slow feature extraction and limited representation accuracy. To address these challenges, we propose NeuroVoxel-LM, a novel framework that integrates Neural Radiance Fields (NeRF) with dynamic resolution voxelization and lightweight meta-embedding. Specifically, we introduce a Dynamic Resolution Multiscale Voxelization (DR-MSV) technique that adaptively adjusts voxel granularity based on geometric and structural complexity, reducing computational cost while preserving reconstruction fidelity. In addition, we propose the Token-level Adaptive Pooling for Lightweight Meta-Embedding (TAP-LME) mechanism, which enhances semantic representation through attention-based weighting and residual fusion. Experimental results demonstrate that DR-MSV significantly improves point cloud feature extraction efficiency and accuracy, while TAP-LME outperforms conventional max-pooling in capturing fine-grained semantics from NeRF weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。