arXiv:2606.00987cs.CVcs.AI2026-06被引 2

提出首个多时相指代分割任务与基准,推动视觉语言模型理解时间变化。

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

论文配图:An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
图 1 · 摘自论文原文
  • 构建自动数据生成与人工审核的流水线,创建21K高质量多时相图文掩码数据集
  • 提出两阶段训练框架,显著提升对时序变化的精准定位能力
  • 适合关注视频理解、时序推理与视觉语言模型的科研人员

大型视觉语言模型(LVLMs)在视觉理解与语言引导定位方面表现强劲,但其多时相视觉推理能力仍待探索。为此,我们提出多时相指代分割(MTRS)新任务,旨在从多时相图像中分割出语言描述的时间变化。MTRS融合传统指代分割与变化检测,要求同时具备时序对应推理、语言定位与像素级掩码预测能力。我们设计了CRAFT-Agent自动化数据构建流程并辅以人工审核,建立首个MTRS基准MTRefSeg-21K,包含21,000个跨多样场景、视角与领域的高质量多时相图像-文本-掩码三元组。对多种基于VLM与LVLM的模型进行基准测试发现,直接推理表现不佳,任务微调效果有限。为此,我们提出改变感知型LVLM框架MTRefSeg-R1,采用两阶段策略:首先在2万张仅视觉的双时相样本上学习通用时序变化感知,再在MTRefSeg-21K上微调以实现细粒度语言引导的时序定位。MTRefSeg-R1显式建模跨时相视觉差异,对齐语言指令与时序变化,并预测目标变化掩码。大量实验表明,该方法优于现有LVLM基线,验证了MTRS的挑战性与潜力。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temporal visual reasoning remains underexplored. To bridge this gap, we introduce \textbf{Multi-temporal Referring Segmentation (MTRS)}, a new task that aims to segment language-described temporal changes from multi-temporal images. MTRS extends conventional referring segmentation and change detection by jointly requiring temporal correspondence reasoning, language grounding, and pixel-level mask prediction. We propose \textbf{CRAFT-Agent}, an automated data construction pipeline with human auditing, and build \textbf{MTRefSeg-21K}, the first MTRS benchmark, containing 21K high-quality multi-temporal image-text-mask triplets across diverse scenes, viewpoints, and domains. Benchmarking a broad set of VLM- and LVLM-based models reveals that direct inference performs poorly, while task-specific fine-tuning remains limited. To address this, we propose \textbf{MTRefSeg-R1}, a change-aware LVLM framework trained with a two-stage strategy. It first learns general temporal-change perception from 20K vision-only bi-temporal samples, and is then fine-tuned on MTRefSeg-21K for fine-grained language-guided temporal localization. MTRefSeg-R1 explicitly models cross-temporal visual differences, aligns language instructions with temporal variations, and predicts referred change masks. Extensive experiments show that MTRefSeg-R1 achieves strong and often superior performance compared with existing LVLM baselines, demonstrating the challenge and potential of MTRS.

多时相指代分割视觉语言模型变化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。