用自监督预训练模型提升分布式声学传感信号识别泛化能力
A Foundation Model for DAS Signal Recognition and Visual Prompt Tuning of the Pre-trained Model for Downstream Tasks
- 基于掩码自编码器构建DAS信号基础模型,通过自监督学习捕捉深层语义特征
- 仅需0.322%参数微调即达96.94%准确率,比全量微调快45%
- 适用于步态识别、管道泄漏等多种下游任务,适合资源受限场景
分布式声学传感(DAS)技术在多个领域应用日益广泛。然而,因感知环境异质性导致的数据分布差异,制约了数据驱动人工智能模型的跨域泛化能力,并面临标注数据不足的问题。为此,本研究提出一种基于掩码自编码器的DAS信号识别基础模型MAEPD,其在包含635,860个样本的数据集上进行预训练,涵盖步态时空信号、周界安防的2D GASF图像、管道泄漏的2D时频图像,以及鲸鱼鸣叫和地震活动等开放数据集信号,通过自监督掩码重建任务提取深层语义特征。采用视觉提示微调(VPT)方法进行下游任务识别,冻结预训练主干参数,仅微调插入Transformer编码层的小规模可学习视觉提示向量。在NVIDIA GeForce RTX 4080 Super平台上的实验表明,以室内步态识别为下游任务,VPT-Deep方法仅需0.322%参数微调即可达到96.94%分类准确率,优于传统全量微调(FFT)方法0.61%,且训练时间减少45%。该模型在管道泄漏检测中也表现出稳健性能,验证了MAEPD作为基础模型的通用性、高效性与可扩展性。该方法为解决DAS领域信号识别模型泛化能力不足提供了新范式。
原文摘要 · Abstract (English)
Distributed Acoustic Sensing (DAS) technology finds growing applications across various domains. However, data distribution disparities due to heterogeneous sensing environments pose challenges for data-driven artificial intelligence (AI) models, limiting cross-domain generalization and facing a shortage of labeled training data. To address these issues, this study proposes a foundational model for DAS signal recognition based on a Masked Autoencoder, named MAEPD. The MAEPD model is pretrained on a dataset of 635,860 samples, encompassing DAS gait spatiotemporal signals, 2D GASF images for perimeter security, 2D time-frequency images for pipeline leakage, and open-dataset signals including whale vocalizations and seismic activities, using a self-supervised mask reconstruction task to capture deep semantic features of DAS signals. Visual Prompt Tuning (VPT) is employed for downstream recognition tasks. This method freezes the pretrained backbone parameters and fine-tunes only a small set of learnable visual prompt vectors inserted into the Transformer encoder layers. Experiments on the NVIDIA GeForce RTX 4080 Super platform validate MAEPD using indoor gait recognition as a downstream task. The VPT-Deep approach achieves a classification accuracy of 96.94% with just 0.322% of parameters fine-tuned, surpassing the traditional Full Fine Tuning (FFT) method by 0.61% and reducing training time by 45%. The model also exhibits robust performance in pipeline leakage detection, confirming the generality, efficiency, and scalability of MAEPD as a foundational model. This approach offers a novel paradigm for addressing the limited generalization of signal recognition models in the DAS domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。