用局部维度分析自监督语音模型的异常,发现对抗样本会持续抬升早期层维度。
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

- 通过分层局部内在维度(LID)捕捉语音表示的局部几何变化
- 低信噪比扰动使LID上升,高信噪比下良性噪声回归正常而对抗样本维持高位
- 可实现无字幕异常检测,适合语音系统监控与安全评估
自监督语音模型(S3Ms)在下游任务中表现优异,但其学习表征在自然和对抗性扰动下的特性仍不清晰。现有研究依赖表征相似性或全局维度,难以揭示局部几何变化。本文提出GRIDS框架,利用WavLM和wav2vec 2.0各层级表征中的局部内在维度(LID),探究扰动如何改变局部几何结构及其与下游自动语音识别(ASR)性能下降的关系。结果表明:所有低信噪比(SNR)扰动均导致LID上升,并在高SNR下出现分化——良性噪声趋于恢复纯净状态,而对抗输入在早期层保持LID升高。我们发现LID升高与错误率(WER)增加同步,且层级LID特征可用于异常检测(AUROC 0.78–1.00),为无需转录文本的S3Ms监控提供了新路径。
原文摘要 · Abstract (English)
Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation similarity or global dimensionality, offering limited visibility into local geometric changes. We ask: how do perturbations deform local geometry, and do these shifts track downstream automatic speech recognition (ASR) degradation? To address this, we present GRIDS, a framework using Local Intrinsic Dimensionality (LID) across layer-wise representations in WavLM and wav2vec 2.0. We find that LID increases for all low signal-to noise ratio (SNR) perturbations and diverges at high SNR: benign noise converges toward the clean profile, while adversarial inputs retain early-layer LID elevation. We show LID elevation co-occurs with increased WER, and that layer-wise LID features enable anomaly detection (AUROC 0.78-1.00), opening the door to transcript-free monitoring in S3Ms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。