arXiv:2608.29183cs.CV2026-08中稿 · BMVC 2026

用基础模型融合几何与外观特征,提升6D位姿估计精度

Foundational feature fusion for conditional flow matching in 6D pose estimation

论文配图:Foundational feature fusion for conditional flow matching in 6D pose estimation
图 1 · 摘自论文原文
  • 利用几何与外观基础模型特征,无需任务专用编码器
  • 跨注意力融合机制使位姿估计准确率显著提升
  • 减少标注依赖和内存开销,适合工业部署

条件流匹配已推动物体6D位姿估计进展,通过逐步去噪和配准物体表示到观测场景,达到当前最优性能。现有方法需在物体-场景重叠上训练特定任务编码器,并依赖简单的特征融合策略解决位姿歧义。本文提出FunFlow6D,一种基于流匹配的新框架,利用几何与外观基础模型的特征进行位姿估计,无需任务专用编码器训练。同时引入基于交叉注意力的融合机制,动态结合几何与外观特征,为流匹配模块提供更丰富的条件。在BOP基准四个数据集上的实验表明,FunFlow6D优于先前最先进方法,同时降低监督需求和内存开销。大量消融实验验证了各组件的有效性。

原文摘要 · Abstract (English)

Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task-specific encoders supervised on object-scene overlap and rely on trivial feature fusion strategies to resolve pose ambiguities. We present FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoder training. We also introduce a cross attention-based fusion mechanism that dynamically combines geometric and appearance features to provide richer conditioning for the flow matching module. Experiments on four datasets from the BOP benchmark show that FunFlow6D outperforms the previous state of the art while reducing supervision requirements and memory overhead. Extensive ablations validate the contribution of each proposed component. Project website: https://tev-fbk.github.io/FunFlow6D/.

6D位姿估计流匹配基础模型特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。