用神经网络让机器人说话更抗干扰,精度高且轻量易部署。
The Talking Robot: Distortion-Robust Acoustic Models for Robot-Robot Communication
- 用端到端联合训练的语音合成与识别模型,专为机器人通信设计。
- 在0dB信噪比下错误率仅8.3%,远超传统方法。
- 总参数仅210万,运行速度低于13毫秒,适合嵌入式设备。
我们提出Artoo,一种用于机器人间通信的端到端学习声学系统,以神经网络替代传统信号处理。系统由轻量级文本转语音(TTS)发送端(118万参数)与基于Conformer的自动语音识别(ASR)接收端(93.8万参数)组成,通过可微通道联合优化。由于机器人通信无需保留音色、语调等副语言特征,系统只关注在信道失真下最大化识别准确率。通过三阶段协同训练流程,TTS发送端学会生成抗干扰的声学编码,在噪声环境下实现0 dB信噪比下8.3%的词错误率(CER)。整个系统仅需210万参数(8.4 MB),在CPU上端到端延迟低于13毫秒,适用于资源受限的机器人平台。
原文摘要 · Abstract (English)
We present Artoo, a learned acoustic communication system for robots that replaces hand-designed signal processing with end-to-end co-trained neural networks. Our system pairs a lightweight text-to-speech (TTS) transmitter (1.18M parameters) with a conformer-based automatic speech recognition (ASR) receiver (938K parameters), jointly optimized through a differentiable channel. Unlike human speech, robot-to-robot communication is paralinguistics-free: the system need not preserve timbre, prosody, or naturalness, only maximize decoding accuracy under channel distortion. Through a three-phase co-training curriculum, the TTS transmitter learns to produce distortion-robust acoustic encodings that surpass the baseline under noise, achieving 8.3% CER at 0 dB SNR. The entire system requires only 2.1M parameters (8.4 MB) and runs in under 13 ms end-to-end on a CPU, making it suitable for deployment on resource-constrained robotic platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。