arXiv:2508.03243cs.CV2025-08ICCV被引 1

用多视角融合解决单视角无法分辨的物体姿态歧义问题

MVTOP: Multi-View Transformer-based Object Pose-Estimation

  • 通过早期融合各视角特征,利用视线几何建模多视角关系
  • 在自建数据集上超越所有单视角与现有多视角方法
  • 无需深度数据,端到端训练,适合复杂场景姿态估计

我们提出MVTOP,一种基于Transformer的多视角刚体物体姿态估计算法。通过早期融合各视角特征,该方法可解决单视角或事后融合单视角结果时无法处理的姿态歧义问题。MVTOP通过从各相机中心发出的视线来建模多视角几何结构。尽管方法假设特定场景下相机内参和相对朝向已知,但这些参数可在每次推理时变化,因而具有高度灵活性。利用视线信息,模型能正确整合多视角信息并预测准确姿态。为验证方法能力,我们构建了一个合成数据集,其姿态仅能通过整体性多视角方法求解,单视角无法完成。MVTOP在该数据集上优于单视角及所有现有多视角方法,并在YCB-V数据集上达到竞争力表现。据我们所知,目前尚无可靠的整体多视角方法可有效解决此类姿态歧义。该模型可端到端训练,无需额外数据(如深度图)。

原文摘要 · Abstract (English)

We present MVTOP, a novel transformer-based method for multi-view rigid object pose estimation. Through an early fusion of the view-specific features, our method can resolve pose ambiguities that would be impossible to solve with a single view or with a post-processing of single-view poses. MVTOP models the multi-view geometry via lines of sight that emanate from the respective camera centers. While the method assumes the camera interior and relative orientations are known for a particular scene, they can vary for each inference. This makes the method versatile. The use of the lines of sight enables MVTOP to correctly predict the correct pose with the merged multi-view information. To show the model's capabilities, we provide a synthetic data set that can only be solved with such holistic multi-view approaches since the poses in the dataset cannot be solved with just one view. Our method outperforms single-view and all existing multi-view approaches on our dataset and achieves competitive results on the YCB-V dataset. To the best of our knowledge, no holistic multi-view method exists that can resolve such pose ambiguities reliably. Our model is end-to-end trainable and does not require any additional data, e.g., depth.

姿态估计多视角Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。