arXiv:2606.26003cs.CL2026-06

构建首个面向阿尔及利亚方言的端到端语音对话系统。

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

论文配图:Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
图 1 · 摘自论文原文
  • 分模块集成语音识别、理解、生成与合成,统一架构支持全语音交互。
  • 在电信领域数据集上,语音识别错误率低,意图识别准确率高。
  • 为低资源方言提供可复现的语音对话系统基线,适合本地化应用研究。

自动语音与语言技术仍严重偏向高资源语言,难以适用于阿尔及利亚方言等低资源语境。该语言面临无标准拼写、频繁法语混用及标注语音资源稀缺等挑战。本文提出一个完整的阿尔及利亚方言语音到语音对话系统。采用模块化流程,整合自动语音识别(ASR)、自然语言理解(NLU)、检索增强生成与文本到语音合成(TTS),并构建电信领域专用的ASR、NLU和TTS数据集,对各组件微调预训练模型。ASR基于Whisper适配,NLU结合Transformer嵌入与任务导向对话框架,TTS则在新收集的方言语料上训练。实验显示各模块性能优异:ASR具备低词错误率,NLU在意图分类与实体识别上得分高,TTS生成语音质量稳定。该系统为阿尔及利亚方言的端到端对话建模提供了可复现基线。

原文摘要 · Abstract (English)

Automatic speech and language technologies are still heavily biased toward high-resource languages, limiting their applicability to dialectal and low-resource settings such as Algerian Dialect. This language presents additional challenges including lack of standardized orthography, frequent codeswitching with French, and scarcity of annotated speech resources. This paper addresses the problem of building a complete speech-to-speech conversational system for Algerian Dialect. We propose a modular pipeline integrating automatic speech recognition, natural language understanding, retrieval-augmented generation, and text-to-speech synthesis within a unified architecture. This work is the continuation of our previous work on Algerian dialectal conversational systems Bechiri and Lanasri [2026], extending it from text-based dialogue modeling to full speech-based interaction. We constructed dedicated datasets for ASR, NLU, and TTS in the telecom domain and fine-tune pretrained models for each component. The ASR system is built on Whisper-based adaptation, while the NLU module combines transformer-based embeddings with a task-oriented dialogue framework. A neural TTS system is trained on a newly collected dialectal corpus to enable spoken response generation. Experimental results show strong performance across all components, including low word error rate for ASR, high intent classification and entity recognition scores for NLU, and stable speech synthesis quality. The proposed system provides a reproducible baseline for end-to-end conversational modeling in Algerian Dialect.

语音对话低资源语言方言识别端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。