Lemon统一处理3D点云与语言,实现高效空间理解。
Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- 将点云与文本融合为单一序列,实现早期跨模态融合。
- 在多个3D任务上达到新基准,支持模型规模扩展。
- 适合需要3D空间推理的机器人、自动驾驶等应用。
将大型多模态模型(LMMs)扩展至3D理解面临独特挑战:点云数据稀疏不规则,现有模型依赖碎片化架构和模态专用编码器,训练过程常不稳定且可扩展性差。我们提出Lemon,一种统一的Transformer架构,通过将3D点云块与语言标记联合处理为单一序列来应对这些挑战。不同于以往依赖模态专用编码器和跨模态对齐模块的方法,该设计实现早期空间-语言融合,消除冗余编码器,提升参数效率,并支持更有效的模型扩展。为应对3D数据复杂性,我们设计了保持空间上下文的结构化分块与标记方案,以及三阶段训练课程,逐步构建从物体级识别到场景级空间推理的能力。Lemon在涵盖物体识别、描述生成到3D场景空间推理的综合性3D理解与推理任务中建立新基准,同时展示出随着模型规模和训练数据增加而持续增强的鲁棒扩展性。本工作为推动现实应用中的3D空间智能提供统一基础。
原文摘要 · Abstract (English)
Scaling large multimodal models (LMMs) to 3D understanding poses unique challenges: point cloud data is sparse and irregular, existing models rely on fragmented architectures with modality-specific encoders, and training pipelines often suffer from instability and poor scalability. We introduce Lemon, a unified transformer architecture that addresses these challenges by jointly processing 3D point cloud patches and language tokens as a single sequence. Unlike prior work that relies on modality-specific encoders and cross-modal alignment modules, this design enables early spatial-linguistic fusion, eliminates redundant encoders, improves parameter efficiency, and supports more effective model scaling. To handle the complexity of 3D data, we develop a structured patchification and tokenization scheme that preserves spatial context, and a three-stage training curriculum that progressively builds capabilities from object-level recognition to scene-level spatial reasoning. Lemon establishes new state-of-the-art performance across comprehensive 3D understanding and reasoning tasks, from object recognition and captioning to spatial reasoning in 3D scenes, while demonstrating robust scaling properties as model size and training data increase. Our work provides a unified foundation for advancing 3D spatial intelligence in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。