无需标注数据,实现真实图像的像素级左右语义理解。
Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images

- 利用3D形状与真实图像数据联合训练,无监督学习左右语义。
- 在未见过的物体类别(如汽车、火车)上仍保持高精度预测。
- 适用于需要对称性理解的视觉任务,如自动驾驶、机器人感知。
尽管已有研究关注3D数据和图像中的反射对称性理解,但真实场景图像的像素级语义左右预测仍面临挑战,主要由于缺乏3D信息、遮挡、物体姿态变化、局部缺失等问题。本文提出一种无监督学习框架,借助近期在3D数据中基于顶点的语义左右理解进展,联合使用3D形状数据集与多样化的真实图像数据,实现单视图图像中像素级语义左右预测。特别地,我们发现一个包含人类及四足动物类形状的中等规模3D形状数据集,结合多样的真实图像数据,足以在完全未见的3D物体类别(如汽车、火车)上实现高质量的语义左右预测。总体而言,该方法在渲染图像与真实图像数据集上,均优于现有最先进方法,在密集像素级语义左右预测上表现更优。
原文摘要 · Abstract (English)
While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images remains challenging, due to certain difficulties including the lack of 3D information, occlusion, object pose variation, partiality, etc. In this work, we propose an unsupervised learning framework to tackle this challenge. Leveraging recent advances in vertex-wise semantic left-right understanding of 3D data, our unsupervised learning method jointly utilises 3D shape and image datasets to infer pixel-wise semantic left-right predictions in single-view images. In particular, we show that a medium-scale 3D shape dataset comprising mainly of human- and quadruped animal-like shapes, combined with diverse in-the-wild image data, are sufficient to achieve high-quality semantic left-right prediction in images, even for entirely unseen 3D object categories, such as cars or trains. Overall, our approach achieves superior performance in dense pixel-wise semantic left-right predictions on both rendered and in-the-wild image datasets when compared to existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。