arXiv:2609.04784cs.CV2026-09

无需训练即可零样本识别物体,提升6D姿态估计精度。

CLON: Cue-Calibrated Linguistic Object Onboarding for Zero-Shot 6D Pose Front-Ends

论文配图:CLON: Cue-Calibrated Linguistic Object Onboarding for Zero-Shot 6D Pose Front-Ends
图 1 · 摘自论文原文
  • 用语言记忆和物体提示权重生成高质量候选框。
  • 在7个BOP数据集上检测与分割准确率分别提升8.1和6.2个百分点。
  • 适合新物体快速接入,无需重新训练模型。

零样本6D姿态估计依赖强大的下游求解器,但前端性能常受限于物体提案:需保留部分可见的正例,同时排除语义上合理的干扰项。本文提出线索校准语言物体接入(CLON)方法,无需针对新物体进行任务特化训练。给定目标物体的渲染模板,CLON构建语言语义记忆用于自上而下的提案生成,并计算物体集合的提示权重用于校准提案评分。语言记忆引导SAM-3生成高召回率提案,提示权重在场景推理前一次性计算并固定不变。在七个BOP-Classic-Core数据集上,相比CNOS和SAM-6D前端,CLON使检测平均精度(AP)提升8.1个百分点,分割AP提升6.2个百分点,下游6D姿态平均召回率(AR)最高提升4.1个百分点。

原文摘要 · Abstract (English)

Zero-shot 6D pose estimation pipelines increasingly rely on strong downstream pose solvers, but their performance is often limited by the front-end: object proposals must preserve partially visible true positives while rejecting semantically plausible distractors. We introduce Cue-Calibrated Linguistic Object Onboarding (CLON), a front-end requiring no task-specific training for new objects. Given rendered templates of the onboarded object set, CLON constructs a linguistic semantic memory for top-down proposal generation and object-set cue weights for calibrated proposal scoring. The linguistic memory guides SAM 3 toward high-recall proposals for onboarded objects, while cue weights are computed once from the onboarded object set before scene inference and kept fixed during online scoring. On seven BOP-Classic-Core datasets, CLON improves detection AP by 8.1 percentage points (pp), segmentation AP by 6.2 pp, and downstream 6D pose AR by up to 4.1 pp over CNOS and SAM-6D front-ends.

6D姿态估计零样本物体检测语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。