利用运动信息提升无监督视频像素匹配精度
Leveraging Motion Information for Better Self-Supervised Video Correspondence Learning
- 设计运动增强模块捕捉物体动态特征
- 多聚类采样策略提升关键运动区域匹配
- 在视频分割与关键点追踪任务上超越现有方法
无监督视频对应学习依赖于在帧间准确关联同一视觉对象的像素,但实现可靠像素匹配仍具挑战。现有方法虽尝试通过特征学习生成唯一像素表示,仍难以实现精确匹配,常出现误匹配,限制其在自监督场景下的效果。为此,本文提出高效的自监督视频对应学习框架MER,旨在从无标签视频中精确提取物体细节。首先,设计专用运动增强引擎,强调捕捉视频中物体的动态运动;其次,引入灵活的跨像素对应信息采样策略(多聚类采样器),使模型更关注运动中重要物体的像素变化。实验表明,该算法在视频对象分割和视频对象关键点追踪等任务上优于当前最先进方法。
原文摘要 · Abstract (English)
Self-supervised video correspondence learning depends on the ability to accurately associate pixels between video frames that correspond to the same visual object. However, achieving reliable pixel matching without supervision remains a major challenge. To address this issue, recent research has focused on feature learning techniques that aim to encode unique pixel representations for matching. Despite these advances, existing methods still struggle to achieve exact pixel correspondences and often suffer from false matches, limiting their effectiveness in self-supervised settings. To this end, we explore an efficient self-supervised Video Correspondence Learning framework (MER) that aims to accurately extract object details from unlabeled videos. First, we design a dedicated Motion Enhancement Engine that emphasizes capturing the dynamic motion of objects in videos. In addition, we introduce a flexible sampling strategy for inter-pixel correspondence information (Multi-Cluster Sampler) that enables the model to pay more attention to the pixel changes of important objects in motion. Through experiments, our algorithm outperforms the state-of-the-art competitors on video correspondence learning tasks such as video object segmentation and video object keypoint tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。