arXiv:2608.03851cs.CV2026-08中稿 · ed

轻量级3D重建模型,融合语义先验与多视角几何推理。

LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

论文配图:LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation
图 1 · 摘自论文原文
  • 用语义描述增强代价体积,结合专家混合机制自适应聚合深度假设。
  • 在ScanNetv2和7-Scenes上实现高质量深度估计与3D重建,效率优于现有方法。
  • 适合机器人、AR等需时空一致4D建模的实时3D感知场景。

实时3D感知对机器人、增强现实及具身智能至关重要。现有多视图立体(MVS)方法依赖几何对应,在无纹理或重复区域表现不佳;单目深度模型虽有强图像级先验,却缺乏多视角几何约束。更重要的是,在机器人与具身操作中,高质量3D几何不仅是静态重建所需,更是学习时序一致4D表示的关键基础。为此,我们提出LiteMVS,一种轻量级多视图深度估计模型,将平面扫掠几何推理与强单目语义和结构先验相结合。核心思想是高效注入来自轻量分割模型和大规模视觉基础模型的高层单目知识。具体而言,LiteMVS在代价体积中引入语义描述符,并采用混合专家(MoE)架构实现深度假设间的自适应几何聚合。此外,从视觉基础模型中蒸馏的几何先验进一步强化单目引导,且不增加推理开销。该设计使LiteMVS不仅提升静态场景下的深度估计与3D重建质量,还为后续时间建模与4D表示学习提供更可靠的几何基础。在ScanNetv2和7-Scenes上的实验表明,LiteMVS在保持竞争力效率的同时,实现了高质量的深度预测与3D重建。

原文摘要 · Abstract (English)

Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods primarily rely on geometric correspondences, which often fail in textureless or repetitive regions, while monocular depth models leverage strong image-level priors but lack robust multi-view geometric constraints. More importantly, in robotics and embodied manipulation scenarios, high-quality 3D geometry is not only essential for static reconstruction, but also serves as a critical foundation for learning temporally consistent 4D representations. To obtain visual representations with stronger structural awareness and greater potential for spatiotemporal extension, we present LiteMVS, a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors. The central idea of LiteMVS is to efficiently inject high-level monocular knowledge, obtained from lightweight segmentation models and large-scale vision foundation models, into a multi-view stereo framework. In particular, LiteMVS enriches the cost volume with semantic descriptors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses. Moreover, geometric priors distilled from vision foundation models further strengthen monocular guidance without increasing inference cost. Through this design, LiteMVS not only improves depth estimation and 3D reconstruction quality in static scenes, but also provides a more reliable geometric foundation for subsequent temporal modeling and 4D representation learning. Experiments on ScanNetv2 and 7-Scenes demonstrate that LiteMVS achieves high-quality depth prediction and 3D reconstruction while maintaining competitive efficiency.

3D重建多视图立体轻量模型语义先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。