用视频生成模型实现任意可描述物体的视频分割,支持罕见和动态概念。
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos

- 通过微调视频扩散模型,将生成目标转为预测掩码潜变量。
- 在罕见物体和动态现象上表现优异,跨域性能提升12个IoU点。
- 适合需要泛化分割能力的研究者,尤其关注视频理解与生成结合场景。
我们提出REM框架,用于对视频中可通过自然语言描述的广泛概念进行分割。该方法利用在互联网规模数据上预训练的视频扩散模型所学习到的通用视觉-语言映射,通过在小规模指代性对象分割数据集上微调实现。核心思路是保留生成模型完整架构,仅将目标从预测噪声改为预测掩码潜变量。由此得到的模型可在仅训练有限类别的情况下,准确分割稀有及未见物体,并能轻松泛化至非物体类动态概念,如烟雾或雨滴。我们在新提出的指代性视频过程分割基准(Ref-VPS)中验证了这一能力。REM在域内数据集(如Ref-DAVIS)上达到当前最优水平,而在域外任务中最高提升12个IoU点,充分体现了生成预训练的优势。我们还证明,视频生成技术的进步会直接提升分割性能。
原文摘要 · Abstract (English)
We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapping learned by video diffusion models on Internet-scale data by fine-tuning them on small-scale Referring Object Segmentation datasets. Our key insight is to preserve the entirety of the generative model's architecture by shifting its objective from predicting noise to predicting mask latents. The resulting model can accurately segment rare and unseen objects, despite only being trained on a limited set of categories. Additionally, it can effortlessly generalize to non-object dynamic concepts, such as smoke or raindrops, as demonstrated in our new benchmark for Referring Video Process Segmentation (Ref-VPS). REM performs on par with the state-of-the-art on in-domain datasets, like Ref-DAVIS, while outperforming them by up to 12 IoU points out-of-domain, leveraging the power of generative pre-training. We also show that advancements in video generation directly improve segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。