利用多视角弱监督提升遮挡状态下多人分割精度
Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
- 通过多视角信息与弱监督机制增强模型对遮挡人体的感知能力
- 在遮挡场景下相比基线模型提升4.20%的分割准确率
- 适合需要精准人体分割的复杂场景应用,如智能监控
多人分割旨在同时完成人体部位分割与实例关联,融合实例级与部件级信息以实现细粒度的人体理解。现有先进方法在公开数据集上表现良好,但在人体重叠区域分割效果显著下降。基于‘从不同视角看重叠人体可能分离’的直觉,本文提出一种利用多视角信息的新型训练框架,以改善遮挡条件下的多人分割性能。方法在训练中引入人体实例的弱监督信号和多视角一致性损失。由于缺乏合适数据集,我们设计半自动标注策略,从多视角RGB+D数据与3D人体骨架生成人体实例分割掩码。实验表明,该方法在遮挡场景下相较基线模型可实现最高4.20%的相对性能提升。
原文摘要 · Abstract (English)
Multi-human parsing is the task of segmenting human body parts while associating each part to the person it belongs to, combining instance-level and part-level information for fine-grained human understanding. In this work, we demonstrate that, while state-of-the-art approaches achieved notable results on public datasets, they struggle considerably in segmenting people with overlapping bodies. From the intuition that overlapping people may appear separated from a different point of view, we propose a novel training framework exploiting multi-view information to improve multi-human parsing models under occlusions. Our method integrates such knowledge during the training process, introducing a novel approach based on weak supervision on human instances and a multi-view consistency loss. Given the lack of suitable datasets in the literature, we propose a semi-automatic annotation strategy to generate human instance segmentation masks from multi-view RGB+D data and 3D human skeletons. The experiments demonstrate that the approach can achieve up to a 4.20\% relative improvement on human parsing over the baseline model in occlusion scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。