arXiv:2608.05137cs.CV2026-08中稿 · ACM MM 2026

动态调度多模态信息,让3D场景理解更精准高效

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

论文配图:SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
图 1 · 摘自论文原文
  • 根据语义需求自动选择最相关模态,避免无关信息干扰
  • 在5个3D理解基准上达到顶尖性能,支持细粒度语义分析
  • 适合需要灵活处理视觉与几何信息的智能体研究者

3D场景理解是具身智能的基础,需融合视觉与几何等异构模态信息进行联合推理。然而,不同任务对模态的相关性差异显著。现有多模态大模型通常采用固定模态组合,忽略查询依赖的模态需求,导致无关模态引入语义噪声,关键模态未能充分使用,造成计算浪费与推理稀释。本文提出SmartMage,一个统一的多模态大模型,通过动态编排异构模态实现语义感知的3D场景理解。具体包括:(1) 基于语义先验、文本-模态对齐与模态质量的语义引导模态自适应路由(SMART)模块,动态选择任务相关模态;(2) 利用模态先验指导专家激活的模态感知门控专家(MAGE)模块,促进多模态推理中的自适应专业化。实证表明,SmartMage在五个3D场景理解基准上达到当前最优表现,并在仅含RGB视频的基准上取得有竞争力结果。在诊断性基准ScanFacet中,任务按细粒度语义类别划分,揭示了不同语义类型偏好的模态组合,进一步验证了SmartMage的有效性。

原文摘要 · Abstract (English)

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies across queries. Existing Multimodal Large Language Models (MLLMs) typically rely on fixed modality combinations, overlooking query-dependent modality needs. Such a rigid design can introduce semantic noise from irrelevant modalities while underutilizing more informative ones, leading to wasted computation and diluted reasoning. To address these challenges, this paper proposes SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic-aware 3D scene understanding. Specifically, SmartMage incorporates: (1) a Semantic-guided Modality Adaptive RouTing (SMART) module that selects task-relevant modalities using semantic priors, text-modality alignment, and modality quality; and (2) a Modality-Aware Gating Expert (MAGE) module that leverages modality priors to guide expert activation, fostering adaptive specialization in multimodal reasoning. Empirically, SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks, and attains competitive results on RGB-only video understanding benchmarks. In our diagnostic benchmark ScanFacet, tasks are divided into fine-grained semantic categories, enabling analysis of modality combinations preferred by each semantic type. The observed modality-semantic patterns provide further evidence of SmartMage's effectiveness. Project page: https://yuecheong.github.io/SmartMage/.

3D理解多模态动态路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。