arXiv:2503.17646cs.SDcs.CV2025-03被引 1

用音频预训练提升振动传感的 crowd 监测精度,减少对标注数据依赖。

Leveraging Audio Representations for Vibration-Based Crowd Monitoring in Stadiums

  • 用公开音频数据无监督预训练,再微调振动数据
  • 实测误差降低5.8倍,显著提升监测精度
  • 适合关注隐私保护与低标注成本的智能场馆应用

体育场馆的观众监控对保障公共安全和提升观赛体验至关重要。现有方法多依赖摄像头和麦克风,易造成干扰并引发隐私担忧。本文提出一种基于地板振动感知的新方法,该方式更具非侵入性。由于体育场这类大型公共场所的振动数据标注困难,训练数据稀缺成为主要挑战。为此,我们提出 ViLA(Vibration Leverage Audio),通过在未标注的跨模态数据上预训练来减少对标注数据的依赖。ViLA 先在音频数据上进行无监督预训练,再用少量领域内振动数据进行微调。利用公开音频数据集(YouTube8M)学习声波特征,并将其迁移到振动信号表示中,有效降低对特定场景振动数据的依赖。真实场景实验表明,使用音频数据预训练的模型相比无预训练模型,误差降低达5.8倍。

原文摘要 · Abstract (English)

Crowd monitoring in sports stadiums is important to enhance public safety and improve the audience experience. Existing approaches mainly rely on cameras and microphones, which can cause significant disturbances and often raise privacy concerns. In this paper, we sense floor vibration, which provides a less disruptive and more non-intrusive way of crowd sensing, to predict crowd behavior. However, since the vibration-based crowd monitoring approach is newly developed, one main challenge is the lack of training data due to sports stadiums being large public spaces with complex physical activities. In this paper, we present ViLA (Vibration Leverage Audio), a vibration-based method that reduces the dependency on labeled data by pre-training with unlabeled cross-modality data. ViLA is first pre-trained on audio data in an unsupervised manner and then fine-tuned with a minimal amount of in-domain vibration data. By leveraging publicly available audio datasets, ViLA learns the wave behaviors from audio and then adapts the representation to vibration, reducing the reliance on domain-specific vibration data. Our real-world experiments demonstrate that pre-training the vibration model using publicly available audio data (YouTube8M) achieved up to a 5.8x error reduction compared to the model without audio pre-training.

振动传感跨模态预训练人群监控隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。