X2SAM让大模型能同时理解图文指令并精准分割图像与视频中的物体。
X2SAM: Any Segmentation in Images and Videos

- 用语言模型+掩码记忆模块实现跨模态语义理解与一致分割
- 支持交互式图文提示,在视频中保持对象跟踪一致性
- 首个统一处理图像视频、兼顾对话与视觉引导的分割模型
多模态大语言模型(MLLM)在图像级视觉理解上表现优异,但对图像和视频的像素级感知仍有限。基础分割模型如SAM系列虽生成高质量掩码,却依赖低级视觉提示,无法原生理解复杂对话指令。现有分割型MLLM通常仅针对图像或视频,且难以在同一界面支持文本与视觉提示。我们提出X2SAM,一个统一的分割型MLLM,将任意分割能力从图像扩展至视频。给定对话指令与视觉提示,X2SAM通过语言模型与掩码记忆模块结合,存储受引导的视觉特征,实现时序一致的视频掩码生成。同一架构支持通用、开放词汇、指代、推理、具身对话生成、交互式及视觉具身分割,适用于图像与视频输入。我们还引入视频视觉定位(V-VGD)基准,评估模型从交互式视觉提示中分割视频目标轨迹的能力。通过异构图像与视频数据集的联合训练策略,X2SAM在视频分割上表现强劲,在图像分割基准上保持竞争力,并保留通用图像与视频对话能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos remains limited. Foundation segmentation models such as the SAM series produce high-quality masks, but they rely on low-level visual prompts and cannot natively interpret complex conversational instructions. Existing segmentation MLLMs narrow this gap, but are usually specialized for either images or videos and rarely support both textual and visual prompts in one interface. We introduce X2SAM, a unified segmentation MLLM that extends any-segmentation capabilities from images to videos. Given conversational instructions and visual prompts, X2SAM couples an LLM with a Mask Memory module that stores guided vision features for temporally consistent video mask generation. The same formulation supports generic, open-vocabulary, referring, reasoning, grounded conversation generation, interactive, and visual grounded segmentation across image and video inputs. We further introduce the Video Visual Grounded (V-VGD) segmentation benchmark, which evaluates whether a model can segment object tracks in videos from interactive visual prompts. With a unified joint training strategy over heterogeneous image and video datasets, X2SAM delivers strong video segmentation performance, remains competitive on image segmentation benchmarks, and preserves general image and video chat ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。