用神经网络替代传统语音传输系统,提升短波通信清晰度。
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
- 用自编码器将语音特征转为连续幅度调制符号,直接传输。
- 在真实短波信道中,语音可懂度显著优于传统模拟与数字系统。
- 兼顾低峰均功率比(<1dB)和抗噪声、多径干扰能力,适合军事通信。
语音压缩常用于移动电话和对讲机等无线通信场景。传统系统将语音编解码器与前向纠错、调制及射频硬件结合使用。本文提出一种自编码器,以神经网络替代多个传统信号处理模块。编码器接收声学特征(短时谱、基频、有无音),生成离散时间但连续取值的正交振幅调制(QAM)符号;通过正交频分复用(OFDM)在高频(HF)无线电通道上传输。解码器将接收到的QAM符号还原为适合合成的声学特征。该自编码器经过训练,在加性高斯噪声和多径信道失真下仍具鲁棒性,同时保持峰均功率比(PAPR)低于1 dB。在仿真与真实短波信道测试中,输出语音可懂度明显超越现有模拟与数字无线电系统,在多种信噪比条件下表现优异。
原文摘要 · Abstract (English)
Speech compression is commonly used to send voice over radio channels in applications such as mobile telephony and two-way push-to-talk (PTT) radio. In classical systems, the speech codec is combined with forward error correction, modulation and radio hardware. In this paper we describe an autoencoder that replaces many of the traditional signal processing elements with a neural network. The encoder takes a vocoder feature set (short term spectrum, pitch, voicing), and produces discrete time, but continuously valued quadrature amplitude modulation (QAM) symbols. We use orthogonal frequency domain multiplexing (OFDM) to send and receive these symbols over high frequency (HF) radio channels. The decoder converts received QAM symbols to vocoder features suitable for synthesis. The autoencoder has been trained to be robust to additive Gaussian noise and multipath channel impairments while simultaneously maintaining a Peak To Average Power Ratio (PAPR) of less than 1 dB. Over simulated and real world HF radio channels we have achieved output speech intelligibility that clearly surpasses existing analog and digital radio systems over a range of SNRs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。