arXiv:2505.21079cs.CV2025-05被引 5

用专家路由机制实现多模态3D场景的自适应融合

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

  • 通过可学习路由选择专家,动态处理不同模态输入
  • 在多个3D基准上表现优于现有方法,提升场景理解精度
  • 适合需要融合图像、点云等多源3D数据的科研与工程应用

近年来,多模态大语言模型(MLLM)在全面理解3D场景方面展现出巨大潜力。然而,现有方法通常仅使用一种或有限的3D模态,导致3D场景表征不完整,影响解释准确性。此外,不同查询任务对模态依赖各异,统一处理所有模态特征可能无法有效捕捉特定上下文。为此,我们提出Uni3D-MoE,一种基于稀疏专家混合(MoE)的3D MLLM,实现自适应多模态融合。该框架集成多种3D模态,包括多视角RGB与深度图、鸟瞰图(BEV)地图、点云和体素表示。核心在于在稀疏MoE架构中引入可学习路由机制,可在令牌级别动态选择合适专家。每个专家根据学习到的模态偏好专精处理特定模态组合,从而灵活协作以满足多样化任务需求。在标准3D场景理解基准及专用数据集上的广泛评估验证了其有效性。

原文摘要 · Abstract (English)

Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typically utilize only one or a limited subset of 3D modalities, resulting in incomplete representations of 3D scenes and reduced interpretive accuracy. Furthermore, different types of queries inherently depend on distinct modalities, indicating that uniform processing of all modality tokens may fail to effectively capture query-specific context. To address these challenges, we propose Uni3D-MoE, a sparse Mixture-of-Experts (MoE)-based 3D MLLM designed to enable adaptive 3D multimodal fusion. Specifically, Uni3D-MoE integrates a comprehensive set of 3D modalities, including multi-view RGB and depth images, bird's-eye-view (BEV) maps, point clouds, and voxel representations. At its core, our framework employs a learnable routing mechanism within the sparse MoE-based large language model, dynamically selecting appropriate experts at the token level. Each expert specializes in processing multimodal tokens based on learned modality preferences, thus facilitating flexible collaboration tailored to diverse task-specific requirements. Extensive evaluations on standard 3D scene understanding benchmarks and specialized datasets demonstrate the efficacy of Uni3D-MoE.

3D理解多模态MoE大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。