将大模型转为专家混合架构,提升3D视觉与动作规划能力
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
- 用专家混合机制改造原有大模型,实现高效3D多模态处理
- 在3D问答和任务规划上表现更优,激活参数更少
- 适合做3D感知、机器人导航等需要空间推理的场景
3D视觉与空间推理对于准确理解三维世界至关重要,尤其相较于基于2D图像的传统视觉推理。由于高质量3D数据获取困难,该领域研究近年才逐步兴起。随着强大大语言模型(LLMs)的发展,近年来出现了面向3D视觉的多模态大模型。然而,多数模型仍以3D数据的视觉编码器为主。本文提出将现有密集激活的大语言模型转化为混合专家(MoE)模型,该结构已在多模态处理中证明有效。除利用其指令遵循能力外,还通过附加一个名为Pose-DiT的扩散头,引入新颖的修正流扩散调度器,实现具身任务规划。在3D问答和任务规划任务上的实验表明,3D-MoE框架在减少激活参数的同时实现了性能提升。
原文摘要 · Abstract (English)
3D vision and spatial reasoning have long been recognized as preferable for accurately perceiving our three-dimensional world, especially when compared with traditional visual reasoning based on 2D images. Due to the difficulties in collecting high-quality 3D data, research in this area has only recently gained momentum. With the advent of powerful large language models (LLMs), multi-modal LLMs for 3D vision have been developed over the past few years. However, most of these models focus primarily on the vision encoder for 3D data. In this paper, we propose converting existing densely activated LLMs into mixture-of-experts (MoE) models, which have proven effective for multi-modal data processing. In addition to leveraging these models' instruction-following capabilities, we further enable embodied task planning by attaching a diffusion head, Pose-DiT, that employs a novel rectified flow diffusion scheduler. Experimental results on 3D question answering and task-planning tasks demonstrate that our 3D-MoE framework achieves improved performance with fewer activated parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。