用听觉感知模型检测人类模仿语音,效果优于人耳。
Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection

- 基于听觉感知设计时频调制表示框架,捕捉语音的时空变化特征。
- 段落级时频调制表示在检测任务中表现超越人类听觉水平。
- 适合语音安全、身份认证领域研究者阅读,尤其关注仿声攻击防御。
人类模仿语音比人工智能生成语音更难检测,因其自然度更高,缺乏明显失真或机械特征。为应对这一挑战,本文提出一种基于听觉感知的时频调制(STM)表示框架,利用伽马音滤波器组(GTFB)和伽玛啁啾滤波器组(GCFB)模拟耳蜗过滤特性,分别建模频率选择性和强度依赖的非对称性。该框架联合捕获语音信号的时序与频谱波动,对应于语谱图随时间的变化及与人类听觉相关的频率轴变化。此外,引入段落级STM表示,通过重叠时间窗分析短时调制模式,实现对语音时变特性的高分辨率建模。实验表明,STM表示可有效检测人类模仿语音,准确率接近人类听觉水平;而段落级STM表示性能更优,超越人类感知能力。结果表明,受听觉机制启发的时频建模在检测模仿类语音攻击方面具有巨大潜力,有助于提升语音认证系统的鲁棒性。
原文摘要 · Abstract (English)
Human-imitated speech poses a greater challenge than AI-generated speech for both human listeners and automatic detection systems. Unlike AI-generated speech, which often contains artifacts, over-smoothed spectra, or robotic cues, imitated speech is produced naturally by humans, thereby preserving a higher degree of naturalness that makes imitation-based speech forgery significantly more challenging to detect using conventional acoustic or cepstral features. To overcome this challenge, this study proposes an auditory perception-based Spectro-Temporal Modulation (STM) representation framework for human-imitated speech detection. The STM representations are derived from two cochlear filterbank models: the Gammatone Filterbank (GTFB), which simulates frequency selectivity and can be regarded as a first approximation of cochlear filtering, and the Gammachirp Filterbank (GCFB), which further models both frequency selectivity and level-dependent asymmetry. These STM representations jointly capture temporal and spectral fluctuations in speech signals, corresponding to changes over time in the spectrogram and variations along the frequency axis related to human auditory perception. We also introduce a Segmental-STM representation to analyze short-term modulation patterns across overlapping time windows, enabling high-resolution modeling of temporal speech variations. Experimental results show that STM representations are effective for human-imitated speech detection, achieving accuracy levels close to those of human listeners. In addition, Segmental-STM representations are more effective, surpassing human perceptual performance. The findings demonstrate that perceptually inspired spectro-temporal modeling is promising for detecting imitation-based speech attacks and improving voice authentication robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。