arXiv:2506.01845eess.AScs.LG2025-06中稿 · Interspeech 2025, …被引 4

让语音单位在设备端实时运行,效率提升一半且误差仅增6.5%

On-device Streaming Discrete Speech Units

  • 压缩注意力窗口与模型规模,实现轻量化部署
  • FLOPs减半,字符错误率仅上升6.5%(ML-SUPERB 1h)
  • 适合资源受限设备的实时语音处理场景

离散语音单元(DSUs)通过聚类自监督语音模型(S3Ms)特征获得,因其丰富的语音信息、高传输效率和与大语言模型的无缝集成,在设备端流式语音应用中具有显著优势。然而,传统方法需完整语音输入且依赖计算量大的S3Ms,难以实用。本文通过缩小注意力窗口并减少模型规模,在保持DSU有效性的前提下,将浮点运算量(FLOPs)降低50%,在ML-SUPERB 1小时数据集上字符错误率(CER)相对增加6.5%。结果表明,DSUs在资源受限环境下的实时语音处理中具有巨大潜力。

原文摘要 · Abstract (English)

Discrete speech units (DSUs) are derived from clustering the features of self-supervised speech models (S3Ms). DSUs offer significant advantages for on-device streaming speech applications due to their rich phonetic information, high transmission efficiency, and seamless integration with large language models. However, conventional DSU-based approaches are impractical as they require full-length speech input and computationally expensive S3Ms. In this work, we reduce both the attention window and the model size while preserving the effectiveness of DSUs. Our results demonstrate that we can reduce floating-point operations (FLOPs) by 50% with only a relative increase of 6.5% in character error rate (CER) on the ML-SUPERB 1h dataset. These findings highlight the potential of DSUs for real-time speech processing in resource-constrained environments.

语音处理离散单元轻量化设备端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。