arXiv:2601.12076cs.CVcs.CL2026-01

解决遥感视频目标分割中记忆质量差的问题,提升复杂场景下的定位精度。

CroBIM-V: Memory-Quality Controlled Remote Sensing Referring Video Object Segmentation

  • 基于运动一致性校准初始记忆,避免结构偏差
  • 动态评估记忆质量,只更新高置信度特征
  • 首个大规模遥感视频指代分割数据集,支持严格因果标注

遥感视频指代目标分割(RS-RVOS)面临目标显著性弱、动态场景中视觉信息严重丢失的挑战,难以维持分割过程中的判别性目标表征。现有研究受限于缺乏大规模专用基准,且模型常因初始记忆构建偏倚导致复杂场景下实例定位不准,以及无差别记忆累积导致遮挡或误分类噪声被编码,引发持续误差传播。本文在数据与方法上双管齐下:首先构建首个大规模基准RS-RVOS Bench,包含111个视频序列、约2.5万帧图像和21.3万条时间指代标注;其标注策略严格遵循因果性,语言描述仅基于首帧目标状态生成。其次提出记忆质量感知的在线指代分割框架MQC-SAM,引入时序运动一致性模块校准初始记忆,利用短时运动轨迹先验纠正结构偏差并建立准确记忆锚点;同时设计解耦注意力式记忆融合机制与动态质量评估,选择性更新高置信度语义特征,过滤不可靠信息,有效防止误差累积与传播。在RS-RVOS Bench上的大量实验表明,MQC-SAM达到当前最优性能。

原文摘要 · Abstract (English)

Remote sensing video referring object segmentation (RS-RVOS) is challenged by weak target saliency and severe visual information truncation in dynamic scenes, making it extremely difficult to maintain discriminative target representations during segmentation. Moreover, progress in this field is hindered by the absence of large-scale dedicated benchmarks, while existing models are often affected by biased initial memory construction that impairs accurate instance localization in complex scenarios, as well as indiscriminate memory accumulation that encodes noise from occlusions or misclassifications, leading to persistent error propagation. This paper advances RS-RVOS research through dual contributions in data and methodology. First, we construct RS-RVOS Bench, the first large-scale benchmark comprising 111 video sequences, about 25,000 frames, and 213,000 temporal referring annotations. Unlike common RVOS benchmarks where many expressions are written with access to the full video context, our dataset adopts a strict causality-aware annotation strategy in which linguistic references are generated solely from the target state in the initial frame. Second, we propose a memory-quality-aware online referring segmentation framework, termed Memory Quality Control with Segment Anything Model (MQC-SAM). MQC-SAM introduces a temporal motion consistency module for initial memory calibration, leveraging short-term motion trajectory priors to correct structural deviations and establish accurate memory anchoring. Furthermore, it incorporates a decoupled attention-based memory integration mechanism with dynamic quality assessment, selectively updating high-confidence semantic features while filtering unreliable information, thereby effectively preventing error accumulation and propagation. Extensive experiments on RS-RVOS Bench demonstrate that MQC-SAM achieves state-of-the-art performance.

遥感分割视频分割记忆控制SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。