提出无标签的韵律对比评估方法,仅用少量样本即可量化语音表征的韵律区分能力。
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
- 基于ABX任务扩展出无需标注的韵律对比评估框架
- 在英、日、汉三种语言中验证了应力、声调和语调的区分能力
- 模型与层的排序在不同条件下稳定,适合低资源场景
自监督语音模型(S3Ms)对音素差异敏感,但其对韵律差异的敏感性尚未被直接衡量。本文引入韵律ABX,将原有的最小对辨识任务扩展至韵律对比评估,仅需少量样本且无需显式标注。我们构建并发布了一个包含英语和日语最小对的数据集,并结合汉语数据集,用于评估英语重音、日语声调、汉语声调的区分能力。实验表明,在多种测试条件下,模型与层的排名具有高度一致性,表明该方法适用于低资源场景。
原文摘要 · Abstract (English)
Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast in S3M representations via minimal pairs. We introduce prosodic ABX, an extension of this framework to evaluate prosodic contrast with only a handful of examples and no explicit labels. Also, we build and release a dataset of English and Japanese minimal pairs and use it along with a Mandarin dataset to evaluate contrast in English stress, Japanese pitch accent, and Mandarin tone. Finally, we show that model and layer rankings are often preserved across several experimental conditions, making it practical for low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。