用仿人耳机制生成语音表示,实现高效语义建模
Representing Speech Through Autoregressive Prediction of Cochlear Tokens
- 分两阶段:先模拟耳蜗生成离散音符,再用自回归模型建模
- 在SUPERB任务上达到顶尖性能,能生成可还原的音频片段
- 适合研究生物启发语音模型与语音生成的应用者
我们提出AuriStream,一种受人类听觉处理层次启发的语音编码模型。第一阶段将原始音频转换为基于人耳耳蜗的时间-频率表示,并从中提取离散的\textbf{耳蜗音符};第二阶段在耳蜗音符序列上应用自回归序列模型。AuriStream能够学习有意义的音素和词级表示,以及最先进的词汇语义。在多样化的下游SUPERB语音任务中表现优异。此外,该模型能生成音频延续,在频谱空间可视化并解码回音频,揭示其预测机制。综上,我们提出一个双阶段语音表示学习框架,推动更类人化、高效处理各类语音任务的模型发展。
原文摘要 · Abstract (English)
We introduce AuriStream, a biologically inspired model for encoding speech via a two-stage framework inspired by the human auditory processing hierarchy. The first stage transforms raw audio into a time-frequency representation based on the human cochlea, from which we extract discrete \textbf{cochlear tokens}. The second stage applies an autoregressive sequence model over the cochlear tokens. AuriStream learns meaningful phoneme and word representations, and state-of-the-art lexical semantics. AuriStream shows competitive performance on diverse downstream SUPERB speech tasks. Complementing AuriStream's strong representational capabilities, it generates continuations of audio which can be visualized in a spectrogram space and decoded back into audio, providing insights into the model's predictions. In summary, we present a two-stage framework for speech representation learning to advance the development of more human-like models that efficiently handle a range of speech-based tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。