解决多视角部分重叠下的物体匹配问题,无需标注即可训练出鲁棒特征网络。
Self-Supervised Partial Cycle-Consistency for Multi-View Matching
- 提出自监督的局部循环一致性机制,适应视角间部分重叠场景。
- 在DIVOTrack上实现4.3%的F1分数提升,尤其在重叠少、人多场景表现突出。
- 适合用于无标注的多摄像头系统,支持大规模场景理解任务。
在多摄像头系统中,跨部分重叠视角匹配物体至关重要,需视图不变特征提取网络。通过循环一致性训练可避免繁琐标注。本文将循环一致性数学形式扩展至处理部分重叠情况,并引入伪掩码引导损失关注部分重叠区域。此外,提出多种互补的循环变体及时间差异场景采样方案,提升自监督设定下的数据输入质量。在挑战性数据集DIVOTrack上的跨摄像头匹配实验表明,相比自监督最新方法,本方法综合贡献使F1得分提升4.3个百分点。改进对训练数据重叠度降低具有鲁棒性,在需在多人中进行少量匹配的复杂场景中表现显著。使用该方法训练的自监督特征网络在多种多摄像头设置下均能有效匹配物体,为大规模多摄像头场景理解提供可能。
原文摘要 · Abstract (English)
Matching objects across partially overlapping camera views is crucial in multi-camera systems and requires a view-invariant feature extraction network. Training such a network with cycle-consistency circumvents the need for labor-intensive labeling. In this paper, we extend the mathematical formulation of cycle-consistency to handle partial overlap. We then introduce a pseudo-mask which directs the training loss to take partial overlap into account. We additionally present several new cycle variants that complement each other and present a time-divergent scene sampling scheme that improves the data input for this self-supervised setting. Cross-camera matching experiments on the challenging DIVOTrack dataset show the merits of our approach. Compared to the self-supervised state-of-the-art, we achieve a 4.3 percentage point higher F1 score with our combined contributions. Our improvements are robust to reduced overlap in the training data, with substantial improvements in challenging scenes that need to make few matches between many people. Self-supervised feature networks trained with our method are effective at matching objects in a range of multi-camera settings, providing opportunities for complex tasks like large-scale multi-camera scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。