arXiv:2603.21478cs.CLcs.LG2026-03被引 2

构建了台湾台语语音意图数据集,助力低资源方言语音技术研究

TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild

  • 采用关键词匹配与多模态线索挖掘,实现低标注成本的数据扩充
  • 覆盖21位老年使用者,共3000条真实场景语音,用于医疗与家居场景
  • 适合研究低资源语言、无文字语言及可扩展数据构建的学者

语音技术虽快速进步,但众多语言仍因资源匮乏而被忽视。本文提出 extbf{TaigiSpeech},一个面向台湾台语(又称台湾闽南语)的真实世界语音意图数据集,该语言为低资源且以口语为主。数据集由21位老年人录制,共计3000个语音片段,适用于医疗与家庭助手等实际应用场景。为应对标注数据稀缺问题,我们探索两种数据挖掘策略:一是基于大模型伪标注的关键词匹配方法(通过中间语言实现),二是利用视听多模态线索的最小文本监督框架。该设计支持对低资源、无文字口语语言的可扩展数据构建。数据集将按CC BY 4.0许可证发布,项目官网及数据集详见https://kwchang.org/taigispeech。

原文摘要 · Abstract (English)

Speech technologies have advanced rapidly and serve diverse populations worldwide. However, many languages remain underrepresented due to limited resources. In this paper, we introduce \textbf{TaigiSpeech}, a real-world speech intent dataset in Taiwanese Taigi (aka Taiwanese Hokkien/Southern Min), which is a low-resource and primarily spoken language. The dataset is collected from older adults, comprising 21 speakers with a total of 3k utterances. It is designed for practical intent detection scenarios, including healthcare and home assistant applications. To address the scarcity of labeled data, we explore two data mining strategies with two levels of supervision: keyword match data mining with LLM pseudo labeling via an intermediate language and an audio-visual framework that leverages multimodal cues with minimal textual supervision. This design enables scalable dataset construction for low-resource and unwritten spoken languages. TaigiSpeech will be released under the CC BY 4.0 license to facilitate broad adoption and research on low-resource and unwritten languages. The project website and the dataset can be found on https://kwchang.org/taigispeech.

语音识别低资源语言多模态数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。