arXiv:2506.00644cs.CLcs.AI2025-06中稿 · INTERSPEECH 2025被引 5

为口吃评估构建基于临床标准的多模态标注数据集

Clinical Annotations for Automatic Stuttering Severity Assessment

  • 采用临床专家标注,基于真实诊疗标准建立多模态标注体系
  • 涵盖口吃发作、继发行为与紧张度评分,提供高可靠测试集
  • 适合语音分析、临床辅助诊断与口吃评估模型研究者

口吃是一种复杂的言语障碍,需专业人员进行有效评估与治疗。本文在现有FluencyBank数据集基础上,引入基于临床标准的新标注方案。为确保标注质量,我们聘请临床专家对数据进行标注,使结果真实反映临床实践。标注为多模态,包含音频视觉特征,用于检测和分类口吃事件、继发行为及紧张度评分。除个体标注外,还提供基于专家共识的高可靠性测试集,可用于评估标注者一致性与机器学习模型性能。实验与分析表明,该任务复杂,需深厚临床知识支持模型训练与评估。

原文摘要 · Abstract (English)

Stuttering is a complex disorder that requires specialized expertise for effective assessment and treatment. This paper presents an effort to enhance the FluencyBank dataset with a new stuttering annotation scheme based on established clinical standards. To achieve high-quality annotations, we hired expert clinicians to label the data, ensuring that the resulting annotations mirror real-world clinical expertise. The annotations are multi-modal, incorporating audiovisual features for the detection and classification of stuttering moments, secondary behaviors, and tension scores. In addition to individual annotations, we additionally provide a test set with highly reliable annotations based on expert consensus for assessing individual annotators and machine learning models. Our experiments and analysis illustrate the complexity of this task that necessitates extensive clinical expertise for valid training and evaluation of stuttering assessment models.

口吃评估多模态标注临床数据语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。