arXiv:2603.21940cs.CL2026-03中稿 · LREC 2026

首个面向突尼斯方言的语音理解数据集,助力低资源语言智能对话发展。

SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding

  • 采集55名母语者说突尼斯方言,覆盖6个应用场景
  • 构建4165条语句、约5小时语音的SLU数据集
  • 提供基准模型,适合低资源语言研究者使用

语音理解(SLU)旨在从用户查询的语音中提取语义信息,是任务导向对话系统的核心。尽管深度神经网络和预训练语言模型推动了SLU的显著进展,但因缺乏资源,仅少数高资源语言受益。本文提出SLURP-TN,通过录制55位母语者在突尼斯方言中表达的句子,手动翻译自六个SLURP领域,构建了一个包含4165条语句、约5小时音频的语音理解数据集。同时,我们开发了若干基于该数据集的自动语音识别与SLU模型。数据集及基线模型已开源:https://huggingface.co/datasets/Elyadata/SLURP-TN。

原文摘要 · Abstract (English)

Spoken Language Understanding (SLU) aims to extract the semantic information from the speech utterance of user queries. It is a core component in a task-oriented dialogue system. With the spectacular progress of deep neural network models and the evolution of pre-trained language models, SLU has obtained significant breakthroughs. However, only a few high-resource languages have taken advantage of this progress due to the absence of SLU resources. In this paper, we seek to mitigate this obstacle by introducing SLURP-TN. This dataset was created by recording 55 native speakers uttering sentences in Tunisian dialect, manually translated from six SLURP domains. The result is an SLU Tunisian dialect dataset that comprises 4165 sentences recorded into around 5 hours of acoustic material. We also develop a number of Automatic Speech Recognition and SLU models exploiting SLUTP-TN. The Dataset and baseline models are available at: https://huggingface.co/datasets/Elyadata/SLURP-TN.

语音理解低资源语言方言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。