用自监督学习减少人群计数标注数据依赖,提升数据效率。
Multi-View Crowd Counting With Self-Supervised Learning
- 通过神经体积渲染构建场景隐式表示,实现跨视角重建。
- 仅用70%训练数据即达当前最优性能,显著降低标注需求。
- 可无缝集成至现有框架,适合数据稀缺场景下的计数任务。
多视角人群计数(MVC)近年来取得显著进展,但多数方法依赖大量标注数据的全监督学习(FSL)。本文提出SSLCounter,一种基于自监督学习(SSL)的新型框架,利用神经体积渲染缓解对大规模标注数据的依赖。该方法学习场景的隐式表示,通过微分神经渲染实现连续几何形状及复杂视图依赖外观的2D投影重建。由于其固有的灵活性,该核心思想可无缝集成至现有框架。大量实验表明,SSLCounter不仅在多个MVC基准上达到顶尖性能,且仅使用70%训练数据即可保持竞争力,展现出卓越的数据效率。
原文摘要 · Abstract (English)
Multi-view counting (MVC) methods have attracted significant research attention and stimulated remarkable progress in recent years. Despite their success, most MVC methods have focused on improving performance by following the fully supervised learning (FSL) paradigm, which often requires large amounts of annotated data. In this work, we propose SSLCounter, a novel self-supervised learning (SSL) framework for MVC that leverages neural volumetric rendering to alleviate the reliance on large-scale annotated datasets. SSLCounter learns an implicit representation w.r.t. the scene, enabling the reconstruction of continuous geometry shape and the complex, view-dependent appearance of their 2D projections via differential neural rendering. Owing to its inherent flexibility, the key idea of our method can be seamlessly integrated into exsiting frameworks. Notably, extensive experiments demonstrate that SSLCounter not only demonstrates state-of-the-art performances but also delivers competitive performance with only using 70% proportion of training data, showcasing its superior data efficiency across multiple MVC benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。