无需训练,用3D模板和图像提案实现零样本2D物体检测分割
MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- 基于3D未见物体的多视角2D模板与输入图像提案匹配
- 联合绝对与相对相似度,提升复杂场景下匹配鲁棒性
- 引入不确定性感知先验,自动调整提案可靠性,适合零样本场景
本文提出MUSE(Model-based Uncertainty-aware Similarity Estimation),一种无需训练的零样本2D物体检测与分割框架。MUSE利用从3D未见物体渲染的2D多视角模板,以及输入查询图像中提取的2D物体提案。在嵌入阶段,融合类别嵌入与补丁嵌入,通过广义均值池化(GeM)对补丁嵌入进行归一化,高效捕捉全局与局部表征。在匹配阶段,采用结合绝对与相对相似度的联合相似度度量,增强在挑战性场景下的匹配鲁棒性。最后,通过不确定性感知的物体先验对相似度得分进行精炼,以调整提案的可靠性。无需任何额外训练或微调,MUSE在BOP Challenge 2025上取得领先表现,包揽Classic Core、H3和Industrial三个赛道第一名。结果表明,MUSE为零样本2D物体检测与分割提供了一种强大且可泛化的框架。
原文摘要 · Abstract (English)
In this work, we introduce MUSE (Model-based Uncertainty-aware Similarity Estimation), a training-free framework designed for model-based zero-shot 2D object detection and segmentation. MUSE leverages 2D multi-view templates rendered from 3D unseen objects and 2D object proposals extracted from input query images. In the embedding stage, it integrates class and patch embeddings, where the patch embeddings are normalized using generalized mean pooling (GeM) to capture both global and local representations efficiently. During the matching stage, MUSE employs a joint similarity metric that combines absolute and relative similarity scores, enhancing the robustness of matching under challenging scenarios. Finally, the similarity score is refined through an uncertainty-aware object prior that adjusts for proposal reliability. Without any additional training or fine-tuning, MUSE achieves state-of-the-art performance on the BOP Challenge 2025, ranking first across the Classic Core, H3, and Industrial tracks. These results demonstrate that MUSE offers a powerful and generalizable framework for zero-shot 2D object detection and segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。