arXiv:2501.15877cs.HCcs.AI2025-01被引 8

构建首个面向印度语的多语言口吃语音数据集,助力技术与理解进步。

Boli: A dataset for understanding stuttering experience and analyzing stuttered speech

  • 采集印度语口吃者朗读与自发语音,含五类口吃标注
  • 涵盖100+参与者元数据与生活影响问卷,覆盖多语言群体
  • 开源数据集,适合语音识别、口吃研究及无障碍技术开发者

随着对多样化高质量口吃语音数据需求的增长,尤其是在印度语背景下,本文推出「Project Boli」——一个面向多语言口吃语音的数据集,旨在推动对口吃人群的理解与相关技术发展。该数据集包含:(a) 匿名化元数据(性别、年龄、国家、母语)及关于口吃对日常生活影响的问卷响应;(b) 每位参与者录制的朗读语音(使用Rainbow Passage)与自发语音(图像描述任务);(c) 对五类口吃现象(阻断、延长、插入音、音节重复、词重复)的详细标注。我们展示了数据收集流程、口吃者经验总结、口吃严重程度评估及数据技术验证。数据集已开放获取,以支持语音技术持续发展。

原文摘要 · Abstract (English)

There is a growing need for diverse, high-quality stuttered speech data, particularly in the context of Indian languages. This paper introduces Project Boli, a multi-lingual stuttered speech dataset designed to advance scientific understanding and technology development for individuals who stutter, particularly in India. The dataset constitutes (a) anonymized metadata (gender, age, country, mother tongue) and responses to a questionnaire about how stuttering affects their daily lives, (b) captures both read speech (using the Rainbow Passage) and spontaneous speech (through image description tasks) for each participant and (c) includes detailed annotations of five stutter types: blocks, prolongations, interjections, sound repetitions and word repetitions. We present a comprehensive analysis of the dataset, including the data collection procedure, experience summarization of people who stutter, severity assessment of stuttering events and technical validation of the collected data. The dataset is released as an open access to further speech technology development.

语音数据口吃研究多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。