arXiv:2508.04721cs.SDcs.AI2025-08被引 10

打造低延迟电信语音助手,实现实时交互与自动化客服。

Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS

  • 采用量化大模型+流式识别+实时语音合成,构建端到端低延迟系统。
  • 四款专用模型实测延迟低于1.0,支持企业级实时部署。
  • 专为电信场景优化,适合呼叫中心自动化与智能应答需求。

我们提出一种面向电信领域的低延迟语音智能代理系统,适用于实时交互式通信场景,支持呼叫中心自动化、智能语音应答(IVR)及AI驱动的客户服务。该系统由NetoAI开发的四种专用模型构成:TSLAM(4-bit量化电信大模型)、T-VEC(电信嵌入模型)、TTE(电信语音识别模型)和T-Synth(电信文本转语音模型)。系统集成流式语音识别(TTE)、对话理解(TSLAM)、基于电信文档的检索增强生成(RAG)以及实时语音合成(T-Synth),显著提升响应速度与领域适配性。为评估系统性能,我们构建了包含500条真人录制电信问题的数据集,源自RFC文档,模拟真实代理查询。实验表明,TSLAM、TTE与T-Synth均达到实时因子(RTF)低于1.0,满足企业级低延迟部署要求。该系统为下一代电信AI提供了基础,可实现自动客服、故障诊断等功能。

原文摘要 · Abstract (English)

We introduce a low-latency telecom AI voice agent pipeline for real-time, interactive telecommunications use, enabling advanced voice AI for call center automation, intelligent IVR (Interactive Voice Response), and AI-driven customer support. The solution is built for telecom, combining four specialized models by NetoAI: TSLAM, a 4-bit quantized Telecom-Specific Large Language Model (LLM); T-VEC, a Telecom-Specific Embedding Model; TTE, a Telecom-Specific Automatic Speech Recognition (ASR) model; and T-Synth, a Telecom-Specific Text-to-Speech (TTS) model. These models enable highly responsive, domain-adapted voice AI agents supporting knowledge-grounded spoken interactions with low latency. The pipeline integrates streaming ASR (TTE), conversational intelligence (TSLAM), retrieval augmented generation (RAG) over telecom documents, and real-time TTS (T-Synth), setting a new benchmark for telecom voice assistants. To evaluate the system, we built a dataset of 500 human-recorded telecom questions from RFCs, simulating real telecom agent queries. This framework allows analysis of latency, domain relevance, and real-time performance across the stack. Results show that TSLAM, TTE, and T-Synth deliver real-time factors (RTF) below 1.0, supporting enterprise, low-latency telecom deployments. These AI agents -- powered by TSLAM, TTE, and T-Synth -- provide a foundation for next-generation telecom AI, enabling automated customer support, diagnostics, and more.

语音助手低延迟电信AI大模型量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。