arXiv:2511.09771cs.CV2025-11中稿 · ICML

仅用一张参考图实现高精度物体6D姿态跟踪与自动重定位

STORM: Segment, Track, and Object Re-Localization from a Single Image

  • 通过分层空间融合注意力机制,支持单图或多图参考条件输入
  • 在LM-O和YCB-Video上实现比基线更高的无标注跟踪准确率
  • 内置验证器可检测漂移并自动重初始化,适合工业部署场景

精确的6D姿态估计与跟踪是物理AI系统的核心能力,但实际部署仍易出错且依赖人工。现有方法常需CAD模型、手动掩码或逐对象适配,在遮挡或快速运动下表现脆弱,缺乏失效识别机制。本文提出STORM,一种基于单张参考图的统一6D跟踪框架,仅需极少人工输入即可提升鲁棒性。STORM结合:(i) 分层空间融合注意力(HSFA),支持单参考或多参考条件输入,可选视觉语言语义条件以解决实例混淆;(ii) 经BCE训练的跟踪验证器,其连续兼容性得分作为能量指标,用于检测漂移并触发自动重初始化。在LM-O和YCB-Video数据集上的实验表明,STORM在无标注条件下优于强基线,能在严重遮挡和快速视角变化下可靠恢复,计算开销极低。

原文摘要 · Abstract (English)

Accurate 6D pose estimation and tracking are core capabilities for physical AI systems, yet real-world deployment remains brittle and labor-intensive. Many pipelines rely on CAD models, manual masking, or per-object adaptation, and still fail under occlusion or fast motion without a principled way to recognize failure. We propose STORM, a unified framework for reference-conditioned 6D tracking that can operate from a single reference image, with minimal manual input and improved robustness. STORM combines: (i) Hierarchical Spatial Fusion Attention (HSFA), a task-driven reference-query fusion architecture that supports both single-reference and multi-reference conditioning and can optionally use vision-language semantic conditioning to resolve instance ambiguities; and (ii) a BCE-trained tracking verifier whose continuous compatibility logit is used as an energy-like score to detect drift and trigger automatic re-initialization. Experiments on LM-O and YCB-Video show that STORM improves annotation-free pose tracking accuracy over strong baselines and recovers reliably from severe occlusions and rapid viewpoint changes with minimal overhead.

6D姿态估计目标跟踪视觉定位工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。