arXiv:2607.17754cs.CVcs.AI2026-07ICRA被引 6

融合视觉与深度信息,提升未知物体在杂乱场景下的分割精度。

DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation

论文配图:DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation
图 1 · 摘自论文原文
  • 采用可变形注意力机制融合RGB与深度数据
  • 在复杂堆叠场景下实例分割性能超越现有方法
  • 专为物流分拣任务设计,适合实际机器人应用

在物流自动化中,对未知物体进行精准分割对于在杂乱环境中实现高效机器人操作至关重要。诸如箱内抓取和货架取货等任务需要鲁棒的感知能力以应对遮挡、物体形状多变及复杂的空间排列。传统基于RGB的方法因依赖纹理而易过度分割,深度方法则主要关注几何特征,常导致欠分割。为此,我们提出DA-Fusion,一种基于可变形注意力的RGB-D融合Transformer,用于未知物体实例分割。该模型有效结合了RGB与深度数据的优势,在杂乱且多层堆叠的物体环境中显著提升了分割精度。同时,我们构建了专为顶部视角下箱内抓取场景设计的物体杂乱箱数据集(OCBD)。大量实验表明,DA-Fusion在多种环境下均优于当前最优方法,特别适用于现实物流任务。

原文摘要 · Abstract (English)

In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelf-picking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB-based methods tend to over-segment objects due to their reliance on texture, while depth-based methods often under-segment by focusing primarily on geometric features. To address these limitations, we propose DA-Fusion, a deformable attention-based RGB-D fusion Transformer designed for unseen object instance segmentation. DA-Fusion effectively combines the strengths of both RGB and depth data, enhancing segmentation accuracy in cluttered and multi-layered object environments. We also introduce the Object Clutter Bin Dataset (OCBD), a benchmark dataset specifically tailored for evaluating bin-picking scenarios in top-down views. Extensive evaluations demonstrate that DA-Fusion outperforms state-of-the-art methods across diverse environments, making it particularly suited for real-world logistics tasks.

实例分割多模态融合物流机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。