用对比学习让图像像素对运动模糊不变,实现无参考视频对齐
Aligning Motion-Blurred Images Using Contrastive Learning on Overcomplete Pixels
- 在自监督训练中对无标签图像施加运动模糊等变换,学习鲁棒像素特征
- 仅用U-Net+新目标,在真实复杂条件下对未见视频帧实现精准对齐
- 过完备像素可同时编码物体身份与相对位置,适合视觉定位任务
我们提出一种新的对比学习目标,用于学习对运动模糊不变的过完备像素级特征。其他不变性(如姿态、光照或天气)可通过在自监督训练中对无标签图像施加相应变换来学习。我们展示,仅用一个简单的U-Net配合该目标,即可生成对未见移动相机拍摄视频的帧具有实用性的局部特征。通过精心设计的玩具示例,我们还表明,过完备像素可编码图像中物体的身份以及相对于这些物体的像素坐标。
原文摘要 · Abstract (English)
We propose a new contrastive objective for learning overcomplete pixel-level features that are invariant to motion blur. Other invariances (e.g., pose, illumination, or weather) can be learned by applying the corresponding transformations on unlabeled images during self-supervised training. We showcase that a simple U-Net trained with our objective can produce local features useful for aligning the frames of an unseen video captured with a moving camera under realistic and challenging conditions. Using a carefully designed toy example, we also show that the overcomplete pixels can encode the identity of objects in an image and the pixel coordinates relative to these objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。