用专家混合架构提升3D几何重建的可扩展性与适应性
MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- 采用专家混合架构,动态分配特征到专用专家
- 在多个基准上达到领先性能,且无需额外计算
- 适合需要高精度3D重建的视觉应用开发者
近期语言与视觉领域的进展表明,模型规模扩大能持续提升各类任务的表现。在3D视觉几何重建中,大规模训练同样证明能学习通用表示。然而,由于几何监督复杂及3D数据多样性,进一步扩展3D模型仍具挑战。为此,我们提出MoRE,一种基于混合专家(MoE)架构的密集3D视觉基础模型,通过动态路由特征至任务专用专家,使其专注于互补的数据方面,从而提升可扩展性与适应性。为增强真实场景下的鲁棒性,MoRE引入基于置信度的深度精炼模块,稳定并优化几何估计。同时,融合密集语义特征与全局对齐的3D骨干表示,实现高保真表面法向预测。MoRE还采用定制损失函数,确保在多种输入和多个几何任务下稳健学习。大量实验表明,MoRE在多个基准上表现领先,并支持高效下游应用而无需额外计算。
原文摘要 · Abstract (English)
Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction, large-scale training has likewise proven effective for learning versatile representations. However, further scaling of 3D models is challenging due to the complexity of geometric supervision and the diversity of 3D data. To overcome these limitations, we propose MoRE, a dense 3D visual foundation model based on a Mixture-of-Experts (MoE) architecture that dynamically routes features to task-specific experts, allowing them to specialize in complementary data aspects and enhance both scalability and adaptability. Aiming to improve robustness under real-world conditions, MoRE incorporates a confidence-based depth refinement module that stabilizes and refines geometric estimation. In addition, it integrates dense semantic features with globally aligned 3D backbone representations for high-fidelity surface normal prediction. MoRE is further optimized with tailored loss functions to ensure robust learning across diverse inputs and multiple geometric tasks. Extensive experiments demonstrate that MoRE achieves state-of-the-art performance across multiple benchmarks and supports effective downstream applications without extra computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。