无监督分离图像中的目标物体,利用运动与3D几何一致性提升重建鲁棒性。
Object Learning and Robust 3D Reconstruction
- 用运动信息作为线索,无监督识别2D图像中的目标物体。
- 通过3D场景几何一致性检测动态异常物体,生成鲁棒掩码。
- 适合研究无监督视觉理解与3D建模的学者参考。
本论文探讨了神经网络在无监督条件下从图像中分割出感兴趣物体的架构设计与训练方法。2D无监督物体分割的核心挑战在于区分前景目标与背景。FlowCapsules利用运动作为2D场景中目标物体的线索。论文后半部分聚焦于3D应用,目标是从输入图像中检测并移除感兴趣物体。在此任务中,我们利用3D场景的几何一致性来检测不一致的动态物体。由此生成的瞬态物体掩码被用于设计鲁棒优化核函数,以提升非严格拍摄条件下的3D建模效果。本研究旨在展示无监督基于物体方法在计算机视觉中的优势,并提出无需监督定义目标物体或前景物体的可能方向。期望激发社区进一步探索图像理解任务中显式物体表示的可能性。
原文摘要 · Abstract (English)
In this thesis we discuss architectural designs and training methods for a neural network to have the ability of dissecting an image into objects of interest without supervision. The main challenge in 2D unsupervised object segmentation is distinguishing between foreground objects of interest and background. FlowCapsules uses motion as a cue for the objects of interest in 2D scenarios. The last part of this thesis focuses on 3D applications where the goal is detecting and removal of the object of interest from the input images. In these tasks, we leverage the geometric consistency of scenes in 3D to detect the inconsistent dynamic objects. Our transient object masks are then used for designing robust optimization kernels to improve 3D modelling in a casual capture setup. One of our goals in this thesis is to show the merits of unsupervised object based approaches in computer vision. Furthermore, we suggest possible directions for defining objects of interest or foreground objects without requiring supervision. Our hope is to motivate and excite the community into further exploring explicit object representations in image understanding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。