提出通用多模态分割框架,支持任意组合视觉模态的高效训练与推理。
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- 构建包含五种模态的大规模预训练数据集ImageNeXt,基于ImageNet扩展而来。
- 在多个数据集上刷新性能纪录,覆盖深度、事件、红外等多类模态场景。
- 适用于多种模态组合,为跨模态分割提供统一可复用的预训练方案。
近年来的研究表明,多模态线索有助于提升语义分割的鲁棒性。然而,针对多种视觉模态的灵活预训练-微调流程仍待探索。本文提出一种新型多模态学习框架OmniSegmentor,具有两项关键创新:1)基于ImageNet构建大规模多模态预训练数据集ImageNeXt,涵盖五种主流视觉模态;2)设计高效的预训练方法,使模型能有效编码不同模态信息。首次实现一个通用的多模态预训练框架,可在任意模态组合下一致增强模型感知能力。OmniSegmentor在多个多模态语义分割数据集上取得新最优结果,包括NYU Depthv2、EventScape、MFNet、DeLiVER、SUNRGBD和KITTI-360。
原文摘要 · Abstract (English)
Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed OmniSegmentor. It has two key innovations: 1) Based on ImageNet, we assemble a large-scale dataset for multi-modal pretraining, called ImageNeXt, which contains five popular visual modalities. 2) We provide an efficient pretraining manner to endow the model with the capacity to encode different modality information in the ImageNeXt. For the first time, we introduce a universal multi-modal pretraining framework that consistently amplifies the model's perceptual capabilities across various scenarios, regardless of the arbitrary combination of the involved modalities. Remarkably, our OmniSegmentor achieves new state-of-the-art records on a wide range of multi-modal semantic segmentation datasets, including NYU Depthv2, EventScape, MFNet, DeLiVER, SUNRGBD, and KITTI-360.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。