arXiv:2606.24464cs.CV2026-06中稿 · ECCV

用3D几何知识提升文本驱动视频分割的准确性与连贯性

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

论文配图:Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation
图 1 · 摘自论文原文
  • 从单张图像预训练中学习几何一致性表示
  • 通过几何感知蒸馏在视频上实现零样本泛化与最佳性能
  • 适合需要强空间理解的视频理解任务研究者

文本驱动的指代视频对象分割(RVOS)旨在根据自然语言描述定位并分割视频中的目标物体。现有模型通常在2D图像或视频数据集上训练,使用简单分割损失,忽视了帧间几何一致性,导致空间理解能力弱。本文提出几何增强的语言引导视频分割框架GeoLaV,分两阶段进行:第一阶段利用单目新视角合成进行单目几何预训练,使模型通过空间对齐在大规模单图数据集上获得几何一致的视觉表征;第二阶段引入几何感知蒸馏,在视频分割数据集上微调,将通用3D先验模型的3D结构知识迁移至分割任务。该过程增强了3D感知能力,提升了时空连贯性与语言对齐效果。大量实验表明,仅使用图像分割数据即可实现显著的零样本泛化;结合几何感知蒸馏后,在多个RVOS基准上达到领先性能。

原文摘要 · Abstract (English)

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive segmentation losses, which overlooks the geometric consistency across frames and leads to weak spatial understanding. In this paper, we propose Geometry-enhanced Language-guided Video segmentation (GeoLaV), a two-stage framework that distills 3D geometric knowledge from images to enhance text-driven video segmentation. In the first stage, we perform monocular geometry pretraining with monocular novel-view synthesis, enabling the model to acquire geometry-consistent visual representations via spatial alignment on large-scale single-image datasets. In the second stage, we introduce geometry-aware distillation and fine-tune the model on video segmentation datasets, transferring 3D structural knowledge from a general 3D prior model. This process reinforces 3D awareness and improves both spatiotemporal coherence and language grounding in segmentation. Extensive experiments show that our method using only image segmentation data already provides notable zero-shot generalization in RVOS. When combined with geometry-aware distillation for fine-tuning on videos, our method achieves state-of-the-art performance across multiple RVOS benchmarks. The code is available at https://github.com/Tony1882880/GeoLaV.

视频分割几何感知零样本语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。