构建三个新RGB-D实例分割数据集,推动细粒度物体识别发展
IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks
- 提出三个基于实例级别的RGB-D分割数据集
- 验证多种基线模型在新数据集上的性能表现
- 设计简单有效的多模态融合方法,适合机器人应用
图像分割是为日常生活提供人类辅助和增强自主性的关键任务。特别是利用视觉与深度信息的RGB-D分割,相比仅使用RGB的方法,有望实现更丰富的场景理解,因而受到越来越多关注。然而,现有研究主要集中在语义分割,导致实例级RGB-D分割数据集稀缺,使当前方法局限于宽泛类别区分,难以捕捉个体物体所需的细粒度细节。为此,我们引入三个全新的、在实例级别区分的RGB-D实例分割基准数据集,具备广泛适用性,支持从室内导航到机器人操作等多种应用。同时,我们在这些基准上对多种基线模型进行了全面评估,揭示其优劣,为未来研究提供方向。最后,我们提出一种简单而有效的RGB-D数据融合方法,大量实验证明该方法有效,为实现更精细的场景理解提供了稳健框架。
原文摘要 · Abstract (English)
Image segmentation is a vital task for providing human assistance and enhancing autonomy in our daily lives. In particular, RGB-D segmentation-leveraging both visual and depth cues-has attracted increasing attention as it promises richer scene understanding than RGB-only methods. However, most existing efforts have primarily focused on semantic segmentation and thus leave a critical gap. There is a relative scarcity of instance-level RGB-D segmentation datasets, which restricts current methods to broad category distinctions rather than fully capturing the fine-grained details required for recognizing individual objects. To bridge this gap, we introduce three RGB-D instance segmentation benchmarks, distinguished at the instance level. These datasets are versatile, supporting a wide range of applications from indoor navigation to robotic manipulation. In addition, we present an extensive evaluation of various baseline models on these benchmarks. This comprehensive analysis identifies both their strengths and shortcomings, guiding future work toward more robust, generalizable solutions. Finally, we propose a simple yet effective method for RGB-D data integration. Extensive evaluations affirm the effectiveness of our approach, offering a robust framework for advancing toward more nuanced scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。