arXiv:2512.08006cs.SDcs.CL2025-12Conference of the …

用服务化架构让语音合成实时使用高质量音素转换。

Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS

  • 将上下文感知音素转换拆成独立服务,降低主引擎负担。
  • 在保持实时性前提下,音素转换准确率显著提升。
  • 适合移动端和离线场景的低延迟语音合成系统。

轻量级实时文本到语音系统对可访问性至关重要。然而,最高效的语音合成模型常依赖轻量级音素转换器,难以应对上下文相关挑战;而具备深层语言理解的先进音素转换器通常计算开销大,无法实现实时性能。本文分析了语音合成中音素转换质量与推理速度的权衡,提出一种实用框架,通过轻量级上下文感知音素转换策略与面向服务的语音合成架构,将复杂模块作为独立服务运行。该设计将重负载的上下文感知组件与核心语音合成引擎解耦,有效突破延迟瓶颈,实现高质量音素转换的实时应用。实验结果表明,所提系统在保持实时响应的同时,显著提升了发音准确性和语言学准确性,适用于离线及终端设备上的语音合成场景。

原文摘要 · Abstract (English)

Lightweight, real-time text-to-speech systems are crucial for accessibility. However, the most efficient TTS models often rely on lightweight phonemizers that struggle with context-dependent challenges. In contrast, more advanced phonemizers with a deeper linguistic understanding typically incur high computational costs, which prevents real-time performance. This paper examines the trade-off between phonemization quality and inference speed in G2P-aided TTS systems, introducing a practical framework to bridge this gap. We propose lightweight strategies for context-aware phonemization and a service-oriented TTS architecture that executes these modules as independent services. This design decouples heavy context-aware components from the core TTS engine, effectively breaking the latency barrier and enabling real-time use of high-quality phonemization models. Experimental results confirm that the proposed system improves pronunciation soundness and linguistic accuracy while maintaining real-time responsiveness, making it well-suited for offline and end-device TTS applications.

语音合成低延迟音素转换服务化架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。