arXiv:2603.29236cs.CV2026-03被引 1

单目相机实时构建三维场景图,语义与几何感知协同提升精度

M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction

  • 多任务融合设计,通过门控上下文增强特征表达
  • 在NYUDv2上语义准确率提升6.6%,深度误差降低9.4%
  • 可直接接入原有单目SLAM系统,适合机器人实时导航

单目相机因其低成本和易部署性,在机器人感知中备受青睐,但从单帧图像实现可靠实时空间理解仍具挑战。尽管近期多任务密集预测模型提升了像素级深度与语义估计能力,但将其稳定应用于单目建图系统仍不简单。本文提出M2H-MX,一种面向单目空间理解的实时多任务感知模型。该模型在保持多尺度特征表示的同时,引入注册门控全局上下文与受控跨任务交互机制,在轻量解码器中实现深度与语义预测相互促进,满足严格延迟要求。其输出可通过紧凑接口直接集成至未修改的单目SLAM流程。我们在密集预测精度与闭环系统性能两方面进行评估。在NYUDv2数据集上,M2H-MX-L达到领先水平,语义mIoU提升6.6%,深度RMSE降低9.4%。在ScanNet上部署于实时单目建图系统时,相较强基线,平均轨迹误差减少60.7%,并生成更清晰的度量-语义地图。结果表明,现代多任务密集预测可被可靠部署于机器人系统的实时单目空间感知中。

原文摘要 · Abstract (English)

Monocular cameras are attractive for robotic perception due to their low cost and ease of deployment, yet achieving reliable real-time spatial understanding from a single image stream remains challenging. While recent multi-task dense prediction models have improved per-pixel depth and semantic estimation, translating these advances into stable monocular mapping systems is still non-trivial. This paper presents M2H-MX, a real-time multi-task perception model for monocular spatial understanding. The model preserves multi-scale feature representations while introducing register-gated global context and controlled cross-task interaction in a lightweight decoder, enabling depth and semantic predictions to reinforce each other under strict latency constraints. Its outputs integrate directly into an unmodified monocular SLAM pipeline through a compact perception-to-mapping interface. We evaluate both dense prediction accuracy and in-the-loop system performance. On NYUDv2, M2H-MX-L achieves state-of-the-art results, improving semantic mIoU by 6.6% and reducing depth RMSE by 9.4% over representative multi-task baselines. When deployed in a real-time monocular mapping system on ScanNet, M2H-MX reduces average trajectory error by 60.7% compared to a strong monocular SLAM baseline while producing cleaner metric-semantic maps. These results demonstrate that modern multi-task dense prediction can be reliably deployed for real-time monocular spatial perception in robotic systems.

单目建图多任务学习语义分割实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。