arXiv:2607.07230cs.CV2026-07中稿 · publication in Mac…

无需标注数据,通过注意力引导的时空聚类实现视频目标分割

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

论文配图:`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation
图 1 · 摘自论文原文
  • 用注意力筛选关键特征片段,结合轻量级时序聚类对齐部分语义
  • 在无约束多目标场景下保持空间精度与时间连贯性,无需光流或合成运动
  • 适配不同分辨率和运动模式,适合资源受限的自监督视频分析

视频目标分割(VOS)是视频理解的基础任务,要求在多帧中精确勾勒并一致追踪物体。虽然有监督方法性能优异,但依赖密集标注数据,成本高且领域覆盖有限。自监督学习可避免人工标注,但现有方法常难以同时保证空间精度与时间连贯性,尤其在复杂多目标场景中表现不佳。许多方法依赖光流、合成运动信号或特定任务预训练,限制了可扩展性和泛化能力。本文提出一种自监督框架——跨时序一致性与聚类(Cross-Temporal Consistency and Clustering),通过结合注意力引导的标记选择与轻量级时序聚类,学习中层、部件感知的表示。该方法不直接操作像素或整物体,而是利用显著性加权的对称一致性目标,在时间维度上对软部件分配进行对齐。框架采用冻结的Transformer主干网络,搭配轻量模块实现自适应标记选择与多偏移时序对齐,支持在不同分辨率和运动模式下的高效扩展。

原文摘要 · Abstract (English)

Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they rely on densely annotated datasets that are costly to obtain and have limited domain coverage. Self-supervised learning offers a promising alternative by removing the need for manual labels; however, existing approaches often struggle to jointly maintain spatial accuracy and temporal coherence, particularly in unconstrained multi-object scenarios. Many rely on optical flow, synthetic motion cues, or task-specific pretraining, limiting scalability and generalisation. We propose a self-supervised framework, Cross-Temporal Consistency and Clustering, that learns mid-level, part-aware representations by combining attention-guided token selection with lightweight temporal clustering. Instead of operating at the pixel or whole-object level, the method aligns soft part assignments across time using a saliency-weighted symmetric consistency objective. The framework leverages a frozen transformer backbone with lightweight modules for adaptive token selection and multi-offset temporal alignment, enabling efficient scaling across resolutions and motion patterns.

自监督视频分割注意力机制时序聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。