为不同机器人定制生成带深度信息的视频分割数据集,提升模型适配性。
Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation
- 基于3D重建生成可配置的RGB-D视频,支持不同机器人传感器布局
- 在MVPd数据集上微调使特定相机位置的分割性能显著提升
- 适合研究机器人视觉、跨平台视频分割的开发者使用
本文提出一种大规模数据生成方法,用于提升不同机器人形态下的类别无关视频分割效果。通过将机器人具身特征(如传感器类型、安装位置、光照条件)纳入数据生成过程,利用3D重建数据(例如HM3DSem)生成可配置的分割视频。由此构建的大型RGB-D视频全景分割数据集(MVPd)可用于基础模型与视频分割模型的广泛基准测试,并支持聚焦于具身性的视频分割研究。实验表明,在MVPd上微调可有效提升基础模型在特定机器人形态(如特定相机位置)上的迁移表现;同时,引入深度图与相机位姿等3D模态能显著提高分割精度与一致性。项目主页见:https://topipari.com/projects/MVPd
原文摘要 · Abstract (English)
This paper presents a method for generating large-scale datasets to improve class-agnostic video segmentation across robots with different form factors. Specifically, we consider the question of whether video segmentation models trained on generic segmentation data could be more effective for particular robot platforms if robot embodiment is factored into the data generation process. To answer this question, a pipeline is formulated for using 3D reconstructions (e.g. from HM3DSem) to generate segmented videos that are configurable based on a robot's embodiment (e.g. sensor type, sensor placement, and illumination source). A resulting massive RGB-D video panoptic segmentation dataset (MVPd) is introduced for extensive benchmarking with foundation and video segmentation models, as well as to support embodiment-focused research in video segmentation. Our experimental findings demonstrate that using MVPd for finetuning can lead to performance improvements when transferring foundation models to certain robot embodiments, such as specific camera placements. These experiments also show that using 3D modalities (depth images and camera pose) can lead to improvements in video segmentation accuracy and consistency. The project webpage is available at https://topipari.com/projects/MVPd
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。