用事件相机实现自监督运动分割,通过去模糊质量优化模型
EV-LayerSegNet: Self-supervised Motion Segmentation using Event Cameras
- 分层建模场景动态,分离学习仿射光流与分割掩码
- 在模拟数据上达到71%的交并比和87%的检测率
- 无需标注即可训练,适合高速运动场景研究者
事件相机是新型类生物传感器,能以更高时间分辨率捕捉运动变化,因其像素对亮度变化异步响应。因此更适合运动相关任务,如运动分割。然而,事件网络的训练仍具挑战性,因获取真实标签成本高、易出错且频率有限。本文提出EV-LayerSegNet,一种基于事件的自监督卷积神经网络,用于运动分割。受场景动态分层表示启发,我们证明可分别学习仿射光流与分割掩码,并用于输入事件的去模糊。去模糊质量被度量并作为自监督学习损失。我们在仅含仿射运动的模拟数据集上训练与测试网络,达到最高71%的交并比(IoU)和87%的检测率。
原文摘要 · Abstract (English)
Event cameras are novel bio-inspired sensors that capture motion dynamics with much higher temporal resolution than traditional cameras, since pixels react asynchronously to brightness changes. They are therefore better suited for tasks involving motion such as motion segmentation. However, training event-based networks still represents a difficult challenge, as obtaining ground truth is very expensive, error-prone and limited in frequency. In this article, we introduce EV-LayerSegNet, a self-supervised CNN for event-based motion segmentation. Inspired by a layered representation of the scene dynamics, we show that it is possible to learn affine optical flow and segmentation masks separately, and use them to deblur the input events. The deblurring quality is then measured and used as self-supervised learning loss. We train and test the network on a simulated dataset with only affine motion, achieving IoU and detection rate up to 71% and 87% respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。