arXiv:2511.05949cs.CV2025-11

零样本匹配2D多边形,提升三维姿态估计与重建精度

Zero-Shot Polygon Matching with Pre-trained Models for Pose Estimation and Polygon Cloud from Challenging Stereo

  • 用预训练模型+几何约束实现零样本多边形匹配
  • 在5个数据集上匹配面积得分68.60%,领先现有方法32%
  • 适合做三维重建、姿态估计且无需额外训练的场景

尽管立体匹配在0维点和1维线段上已成熟,但2维多边形的对应关系建立仍面临视差不连续、尺度变化、训练依赖和泛化能力差等挑战,限制了姿态估计和3D重建等下游任务。为此,我们首次提出基于预训练模型的零样本多边形匹配框架Z(PM)²,通过可插拔模块融合学习特征与手工几何约束,将匹配从0D/1D扩展至2D多边形。该流程包含三阶段:首先,检测器利用预训练的Segment Anything模型将分割掩码转为含几何与纹理信息的图结构多边形;其次,全局匹配器采用双向金字塔与多几何约束应对视角变化;最后,局部匹配器通过局部-整体双分图优化解决视差不连续与拓扑不一致问题。此外,我们设计了多边形匹配引导的姿态估计方法,获得分布均匀、冗余低的同源点,并首次提出多边形云概念及最优表面生成方法,构建出结构完整、语义丰富的3D表征。因无直接可比的立体图像多边形匹配方法,我们选取相近的最先进方法作为基线。在五个挑战性数据集(ISPRS, KITTI, ScanNet, SceneFlow, DTU)上的大量实验表明,Z(PM)²达到68.60%的匹配面积得分,较MESA提升约32%,在面积级姿态估计中排名第一,具备良好速度与强零样本泛化能力,无需任何训练。

原文摘要 · Abstract (English)

While stereo matching has achieved maturity for 0D point and 1D line primitives, establishing correspondences for 2D polygons remains largely unexplored due to challenges including disparity discontinuity, scale variation, training dependency, and poor generalization, limiting downstream tasks such as pose estimation and 3D reconstruction. To address these issues, we are the first to propose a Zero-shot Polygon Matching paradigm with Pre-trained Models (i.e., Z(PM)2), which combines learned features and handcrafted geometric constraints through plug-and-play modules, extending matching from 0D/1D primitives to 2D polygons. The pipeline comprises three core stages: Firstly, detector leverages the pre-trained segment anything model to vectorize segmentation masks into graph-structured polygons integrating geometry and texture; Secondly, global matcher uses bidirectional-pyramid and multi-geometric constraints to handle viewpoint variation; Thirdly, local matcher leverages local-holistic bipartite graph optimization to resolve disparity discontinuity and topological inconsistency. Moreover, we develop polygon-matching-guided pose estimation using correspondences to obtain well-distributed, low-redundancy homologous points, and pioneer the polygon cloud concept with an optimal surface generation method, producing structurally complete and semantically rich 3D representations beyond point and line clouds. Since no polygon matching methods from stereo imagery are available for direct comparison, we selected state-of-the-art (SoTA) methods close to this task as baselines. Extensive experiments on five challenging datasets (ISPRS, KITTI, ScanNet, SceneFlow, DTU) show Z(PM)2 achieves a 68.60% matching area score, outperforming MESA by approximately 32% and ranking first in area-level pose estimation, with competitive speed and strong zero-shot generalization without any training requirement.

立体匹配多边形云零样本姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。