提出新视频分割任务,让模型理解物体状态变化。
M$^3$-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
- 引入物体'相变'概念,构建多阶段多场景分割数据集
- 479段高清视频,覆盖10种日常场景,标注物体状态变化
- 现有方法在相变场景表现差,新模型通过反向优化提升性能
智能机器人需在多样环境中与物体交互,而物体的外观和状态常随物性发生复杂变化,如相变。但视觉领域长期忽视此类动态对象的分割问题。为此,本文首次引入'相'的概念,基于视觉特征与形态变化潜力对真实世界物体进行分类,并构建新的基准M$^3$-VOS(Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation),包含479个高分辨率视频,覆盖超过10种日常生活场景,提供密集实例掩码标注,精确捕捉物体相态及其转变过程。我们在该数据集上评估了主流方法,发现当前基于外观的方法在处理相变时仍有显著不足;且正向熵增过程的预测性能可通过反向熵减过程改善。据此提出ReVOS,一种可插拔的反向精炼模型,显著提升性能。数据与代码将公开于https://zixuan-chen.github.io/M-cube-VOS.github.io/。
原文摘要 · Abstract (English)
Intelligent robots need to interact with diverse objects across various environments. The appearance and state of objects frequently undergo complex transformations depending on the object properties, e.g., phase transitions. However, in the vision community, segmenting dynamic objects with phase transitions is overlooked. In light of this, we introduce the concept of phase in segmentation, which categorizes real-world objects based on their visual characteristics and potential morphological and appearance changes. Then, we present a new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M$^3$-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10 distinct everyday scenarios. It provides dense instance mask annotations that capture both object phases and their transitions. We evaluate state-of-the-art methods on M$^3$-VOS, yielding several key insights. Notably, current appearance-based approaches show significant room for improvement when handling objects with phase transitions. The inherent changes in disorder suggest that the predictive performance of the forward entropy-increasing process can be improved through a reverse entropy-reducing process. These findings lead us to propose ReVOS, a new plug-andplay model that improves its performance by reversal refinement. Our data and code will be publicly available at https://zixuan-chen.github.io/M-cube-VOS.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。