arXiv:2603.26939cs.SDcs.CL2026-03被引 1

用多语言数据训练,实现跨语言流畅性检测。

Multilingual Stutter Event Detection for English, German, and Mandarin Speech

  • 基于英德汉三语多语料库训练多标签检测模型。
  • 跨语言检测性能达现有水平,部分场景更优。
  • 适合语音分析、言语障碍研究者使用。

本文提出一种基于多语种、多语料库的多标签口吃事件检测系统,涵盖英语、德语和汉语三种语言及四个语料库。通过利用三语言标注的口吃数据,模型捕捉到口吃在不同语言中的共性特征,实现跨语言上下文下的鲁棒检测。实验结果表明,多语言训练在性能上可媲美甚至超过以往系统,在部分情况下表现更优。这些发现表明口吃具有跨语言一致性,支持构建语言无关的检测系统。本研究验证了使用多语言数据提升自动口吃检测泛化能力与可靠性的可行性与优势。

原文摘要 · Abstract (English)

This paper presents a multi-label stuttering detection system trained on multi-corpus, multilingual data in English, German, and Mandarin.By leveraging annotated stuttering data from three languages and four corpora, the model captures language-independent characteristics of stuttering, enabling robust detection across linguistic contexts. Experimental results demonstrate that multilingual training achieves performance comparable to and, in some cases, even exceeds that of previous systems. These findings suggest that stuttering exhibits cross-linguistic consistency, which supports the development of language-agnostic detection systems. Our work demonstrates the feasibility and advantages of using multilingual data to improve generalizability and reliability in automated stuttering detection.

口吃检测多语言语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。