arXiv:2605.12774cs.CV2026-05

统一动态环境下的单目位姿估计,兼顾鲁棒性与精度。

WildPose: A Unified Framework for Robust Pose Estimation in the Wild

论文配图:WildPose: A Unified Framework for Robust Pose Estimation in the Wild
图 1 · 摘自论文原文
  • 融合前馈模型感知与可微分优化,构建3D感知更新机制。
  • 在动态(Wild-SLAM)、静态(TUM)和低运动(Sintel)数据集上均领先。
  • 适合需要高鲁棒性的实时定位与建图场景。

在动态环境中估计相机位姿是关键挑战,因多数视觉SLAM与SfM方法假设场景静止。现有动态感知方法多不统一:基于语义的方法脆弱,逐序列优化对短序列失效,其他学习模型在纯静态场景中性能下降。本文提出WildPose,一种统一的单目位姿估计框架,在动态环境中保持鲁棒性的同时,在静态与低自身运动数据集上达到顶尖性能。核心思想是结合现代3D视觉中两大强大范式:前馈模型的丰富感知前端与可微分束调整(BA)的端到端优化。通过基于冻结预训练MASt3R特征主干的3D感知更新算子,以及利用同一主干的多层级3D感知特征的高容量运动掩码检测器实现。大量实验表明,WildPose在动态(Wild-SLAM、Bonn)、静态(TUM、7-Scenes)和低自身运动(Sintel)基准上持续优于现有方法。

原文摘要 · Abstract (English)

Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-aware methods exist, they are often not unified: semantic-based approaches are brittle, per-sequence optimization methods fail on short sequences, and other learned models may degrade on static-only scenes. We present WildPose, a unified monocular pose estimation framework that is robust in dynamic environments while maintaining state-of-the-art performance on static and low-ego-motion datasets. Our key insight is to connect two powerful paradigms in modern 3D vision: the rich perceptual frontend of feedforward models and the end-to-end optimization of differentiable bundle adjustment (BA). We achieve this with a 3D-aware update operator built on a frozen, pre-trained MASt3R feature backbone, together with a high-capacity motion mask detector that uses multi-level 3D-aware features from the same backbone. Extensive experiments show WildPose consistently outperforms prior methods across dynamic (Wild-SLAM, Bonn), static (TUM, 7-Scenes), and low-ego-motion (Sintel) benchmarks.

位姿估计动态环境单目3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。