无需ASR,用少量成人模板评估儿童发音,但效果受年龄影响大。
Towards few-shot isolated word reading assessment
- 用自监督模型中间层编码语音,对比儿童与成人模板。
- 儿童语音识别准确率显著低于成人,即使使用儿童模板也存在差距。
- 适用于低资源环境下的儿童发音评估,需关注年龄差异影响。
我们探索了一种在低资源环境下进行孤立词发音评估的无自动语音识别(ASR)方法。该少样本方法将儿童语音输入与少量成人提供的参考模板进行比较。输入和模板均通过大规模自监督学习(SSL)模型的中间层进行编码。基于南非荷兰语(Afrikaans)儿童语音基准数据集,我们研究了离散化SSL特征及模板的巴氏平均等设计选项。理想化实验显示,成人语音表现良好,但儿童语音输入即使使用儿童模板,性能仍明显下降。尽管SSL表示在低资源语音任务中表现优异,本工作揭示了其在少样本分类系统中处理儿童数据时的局限性。
原文摘要 · Abstract (English)
We explore an ASR-free method for isolated word reading assessment in low-resource settings. Our few-shot approach compares input child speech to a small set of adult-provided reference templates. Inputs and templates are encoded using intermediate layers from large self-supervised learned (SSL) models. Using an Afrikaans child speech benchmark, we investigate design options such as discretising SSL features and barycentre averaging of the templates. Idealised experiments show reasonable performance for adults, but a substantial drop for child speech input, even with child templates. Despite the success of employing SSL representations in low-resource speech tasks, our work highlights the limitations of SSL representations for processing child data when used in a few-shot classification system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。