arXiv:2412.04117cs.CV2024-12被引 3

无需标注数据,让行人检测模型跨摄像头布局自适应

MVUDA: Unsupervised Domain Adaptation for Multi-view Pedestrian Detection

  • 用伪标签+均值教师框架实现无监督跨视角迁移
  • 在多视角数据集间迁移时准确率超越现有方法
  • 适合部署于不同摄像头配置的实时监控系统

针对多摄像头环境下行人检测模型在训练与测试场景不一致导致性能下降的问题,本文提出一种无监督域自适应方法(MVUDA),可在不依赖额外标注数据的前提下,将模型适配到新的摄像头布局。该方法基于均值教师自训练框架,设计了一种专用于多视角行人检测的伪标签生成策略。实验表明,该方法在多个基准上取得当前最优表现,包括从MultiviewX到Wildtrack的跨域迁移任务。相比以往方法,无需外部单目标注数据,显著降低对标注数据的依赖。大量评估验证了方法的有效性与关键设计的合理性。本工作提升了多视角行人检测系统的实际部署能力,并为后续研究提供了强效的无监督域自适应基线。

原文摘要 · Abstract (English)

We address multi-view pedestrian detection in a setting where labeled data is collected using a multi-camera setup different from the one used for testing. While recent multi-view pedestrian detectors perform well on the camera rig used for training, their performance declines when applied to a different setup. To facilitate seamless deployment across varied camera rigs, we propose an unsupervised domain adaptation (UDA) method that adapts the model to new rigs without requiring additional labeled data. Specifically, we leverage the mean teacher self-training framework with a novel pseudo-labeling technique tailored to multi-view pedestrian detection. This method achieves state-of-the-art performance on multiple benchmarks, including MultiviewX$\rightarrow$Wildtrack. Unlike previous methods, our approach eliminates the need for external labeled monocular datasets, thereby reducing reliance on labeled data. Extensive evaluations demonstrate the effectiveness of our method and validate key design choices. By enabling robust adaptation across camera setups, our work enhances the practicality of multi-view pedestrian detectors and establishes a strong UDA baseline for future research.

行人检测无监督学习域自适应多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。