针对复杂视频目标分割难题,VOS-Agent通过智能路由实现精准分割。
VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)

- 根据目标特性动态激活视觉跟踪或语义理解专用模块
- 在MOSEv2测试集上取得69.82%的J&F得分,排名第一
- 适合处理小目标、消失重现及依赖语义的目标分割
复杂视频对象分割需在严重遮挡、消失与重出现场景下保持目标传播的鲁棒性。尽管SAM3具备强大的可提示掩码传播能力,但统一推理路径对缺乏视觉证据的小目标以及依赖显式属性的语义主导目标仍不可靠。为此,我们提出VOS-Agent,一个协作框架:以SAM3作为共享密集分割模块,并根据目标特征条件性激活专用代理。目标感知与路由代理将每个序列分配至常规、小目标或语义主导路径。小目标由基于置信度感知框提示的视觉跟踪代理支持,语义主导目标则由基于多模态大模型的语义代理通过描述引导定位与候选验证处理。在MOSEv2测试集上,VOS-Agent在官方$/mathcal{J} ext{ extbackslash}& ext{ extbackslash} ext{ extbackslash}dot{ ext{ extbackslash}mathcal{F}}$指标上达到69.82%,并在ECCV 2026第八届LSVOS挑战赛的MOSEv2赛道中排名第一。
原文摘要 · Abstract (English)
Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides strong promptable mask propagation, a uniform inference path remains unreliable for tiny targets with insufficient visual evidence and semantic-dominated targets whose identities depend on explicit attributes. To this end, we present VOS-Agent, a collaborative framework that retains SAM3 as the shared dense segmentation module and conditionally activates specialized agents according to target characteristics. A Target Perception and Routing Agent assigns each sequence to a regular, tiny, or semantic-dominated route. Tiny targets are supported by a Visual Tracking Agent through confidence-aware box prompts, while semantic-dominated targets are handled by an MLLM-based Semantic Agent through description-guided localization and candidate verification. On the MOSEv2 test set, VOS-Agent achieves 69.82% on the official $\mathcal{J}\&\dot{\mathcal{F}}$ metric and ranks first in the MOSEv2 Track of the 8th LSVOS Challenge at ECCV 2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。