arXiv:2512.08430cs.CVcs.RO2025-12中稿 · WACV 2026被引 1

纯深度图实现高精度6D位姿估计,适用于工业密集抓取场景。

SDT-6D: Fully Sparse Depth-Transformer for Staged End-to-End 6D Pose Estimation in Industrial Multi-View Bin Picking

  • 基于多视角深度图构建稀疏体素场,分阶段生成注意力先验
  • 在IPD和MV-YCB数据集上达到领先性能,支持多目标同时预测
  • 适合需要高精度位姿估计的工业机器人抓取任务

在密集堆放的工业料箱抓取环境中准确恢复6D位姿仍面临巨大挑战,主要由于遮挡、反光和无纹理表面。本文提出一种全稀疏的纯深度图6D位姿估计方法,将多视角深度图融合为细粒度3D点云或稀疏截断有符号距离场(TSDF)。核心是分阶段热力图机制,生成跨分辨率的场景自适应注意力先验,引导计算聚焦前景区域,使高分辨率下的内存需求可行。同时提出密度感知稀疏变换器模块,动态关注自身遮挡及3D数据非均匀分布。尽管稀疏3D方法在远距离感知中表现优异,其在近距离机器人应用中的潜力尚未被充分挖掘。本框架完全基于稀疏处理,支持高分辨率体素表示以捕捉关键几何细节。方法整体处理场景,通过新颖的逐体素投票策略预测6D位姿,可同时预测任意数量目标物体的位姿。在新发布的IPD和MV-YCB多视角数据集上验证,展示了在高度杂乱的工业与家庭料箱抓取场景中的竞争力。

原文摘要 · Abstract (English)

Accurately recovering 6D poses in densely packed industrial bin-picking environments remain a serious challenge, owing to occlusions, reflections, and textureless parts. We introduce a holistic depth-only 6D pose estimation approach that fuses multi-view depth maps into either a fine-grained 3D point cloud in its vanilla version, or a sparse Truncated Signed Distance Field (TSDF). At the core of our framework lies a staged heatmap mechanism that yields scene-adaptive attention priors across different resolutions, steering computation toward foreground regions, thus keeping memory requirements at high resolutions feasible. Along, we propose a density-aware sparse transformer block that dynamically attends to (self-) occlusions and the non-uniform distribution of 3D data. While sparse 3D approaches has proven effective for long-range perception, its potential in close-range robotic applications remains underexplored. Our framework operates fully sparse, enabling high-resolution volumetric representations to capture fine geometric details crucial for accurate pose estimation in clutter. Our method processes the entire scene integrally, predicting the 6D pose via a novel per-voxel voting strategy, allowing simultaneous pose predictions for an arbitrary number of target objects. We validate our method on the recently published IPD and MV-YCB multi-view datasets, demonstrating competitive performance in heavily cluttered industrial and household bin picking scenarios.

6D位姿深度估计机器人抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。