arXiv:2607.06620cs.CVcs.AI2026-07中稿 · ECCV

仅用稀疏图像实现3D空间推理,突破多模态模型的几何理解瓶颈。

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

论文配图:SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
图 1 · 摘自论文原文
  • 通过自适应时空采样构建几何感知图,保留场景拓扑结构
  • 引入指令-位姿感知路由的专家混合机制,提升跨模态融合效率
  • 在多个3D推理任务中领先,尤其在路径规划与方向判断上提升超50%

当前多模态大模型难以弥合2D语义理解与3D空间几何之间的表征鸿沟。现有3D感知模型或依赖昂贵的3D数据,或仅使用RGB输入并采用启发式采样与单一浅层融合,分别破坏了时空连续性并引发多模态任务间的冲突。为此,我们提出SpaR3D-MoE,一种端到端框架,仅通过稀疏RGB输入即可赋予大模型自适应的几何感知空间推理能力。首先,设计自适应时空流形采样机制,构建几何感知时空图以提取关键帧,有效减少序列冗余同时保持场景拓扑连通性。其次,引入异构几何归纳的专家混合(MoE)架构,由指令-位姿感知路由器动态分配多模态令牌至专用专家,缓解单体融合带来的跨模态冲突。在VSI-Bench、ScanQA和SQA3D上的大量实验表明,本方法达到最先进性能:在VSI-Bench上平均得分63.5,超越最强基线7.8分;在路径规划与相对方向任务中分别实现35.4%和51.4%的相对提升。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatial geometry. Existing 3D-aware models either rely on costly 3D-specific data or utilize RGB-only inputs with heuristic sampling and monolithic, shallow fusion, which respectively disrupt essential spatiotemporal connectivity and induce modality contention across diverse spatial tasks. To overcome these bottlenecks, we introduce SpaR3D-MoE, an end-to-end framework that enables adaptive spatial reasoning by equipping MLLMs with geometry-aware capabilities from only sparse RGB inputs. First, we propose an adaptive spatiotemporal manifold sampling mechanism that constructs a geometry-aware spatiotemporal graph to extract informative keyframes, effectively mitigating sequence redundancy while preserving the scene's topological connectivity. Second, we introduce the heterogeneous geometry-inductive Mixture-of-Experts driven by an instruction-pose aware router, which adaptively routes multimodal tokens to specialized experts, resolving the cross-modal contention inherent in monolithic fusion. Extensive experiments on VSI-Bench, ScanQA, and SQA3D demonstrate that our method achieves state-of-the-art performance. Notably, SpaR3D-MoE achieves the highest average score of 63.5 on VSI-Bench, outperforming the strongest baseline by 7.8 absolute points, alongside relative improvements of 35.4% and 51.4% in Route Plan and Relative Direction tasks, respectively.

3D推理多模态专家混合几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。