通过混合生成清晰中间帧,提升小重叠图像的位姿估计精度
PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
- 融合视频插值与位姿引导的新视角合成,生成更清晰中间帧
- 在无重叠或极小重叠场景下,位姿估计准确率显著优于现有方法
- 新设计的特征匹配选择器快速筛选适合位姿估计的帧,效率更高
从稀疏重叠图像对中进行配对相机位姿估计仍是三维视觉中的关键难题。现有方法在重叠区域极小或无重叠时表现不佳。近期方法尝试通过视频插值生成中间帧,并利用自一致性分数选择关键帧,但生成帧常模糊,且选择策略缓慢且未直接对齐位姿估计目标。为此,本文提出混合视频生成(HVG),通过耦合视频插值模型与位姿条件的新视角合成模型,生成更清晰的中间帧;同时提出基于特征对应关系的特征匹配选择器(FMS),从合成结果中选出适合作位姿估计的帧。在Cambridge Landmarks、ScanNet、DL3DV-10K和NAVI数据集上的大量实验表明,相比现有最先进方法,PoseCrafter在小重叠或无重叠情况下均显著提升位姿估计性能。
原文摘要 · Abstract (English)
Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video interpolation and selecting key frames via a self-consistency score. However, the generated frames are often blurry due to small overlap inputs, and the selection strategies are slow and not explicitly aligned with pose estimation. To solve these cases, we propose Hybrid Video Generation (HVG) to synthesize clearer intermediate frames by coupling a video interpolation model with a pose-conditioned novel view synthesis model, where we also propose a Feature Matching Selector (FMS) based on feature correspondence to select intermediate frames appropriate for pose estimation from the synthesized results. Extensive experiments on Cambridge Landmarks, ScanNet, DL3DV-10K, and NAVI demonstrate that, compared to existing SOTA methods, PoseCrafter can obviously enhance the pose estimation performances, especially on examples with small or no overlap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。