融合自监督与人工特征,更好建模口语评分的等级顺序与不均等间隔。
An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
- 结合声学与文本自监督特征,引入人工指标增强表达能力。
- 在TEEMI数据集上优于强基线,对未见题目泛化能力强。
- 提出多间隔有序损失,同时捕捉分数等级与非均匀间隔特性。
近年来,自动口语评估(ASA)研究受益于自监督学习(SSL)表示,其无需预设特征工程即可捕捉非母语口语中的丰富声学与语言模式。然而,基于语音的SSL模型仅关注声学特征而忽略语言内容,基于文本的SSL模型依赖语音识别输出,无法编码语调细节。此外,多数现有方法将语言水平视为无序类别,忽视了其等级结构及水平标签间的非均匀间隔。为解决上述问题,本文提出一种新方法:结合自监督学习与人工设计的指示特征,采用新颖建模范式,并引入多间隔有序损失函数,联合建模评分等级性与非均匀间隔。在TEEMI语料库上的大量实验表明,该方法持续优于强基线,且对未见过的题目具有良好的泛化能力。
原文摘要 · Abstract (English)
A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of feature curation. However, speech-based SSL models capture acoustic-related traits but overlook linguistic content, while text-based SSL models rely on ASR output and fail to encode prosodic nuances. Moreover, most prior arts treat proficiency levels as nominal classes, ignoring their ordinal structure and non-uniform intervals between proficiency labels. To address these limitations, we propose an effective ASA approach combining SSL with handcrafted indicator features via a novel modeling paradigm. We further introduce a multi-margin ordinal loss that jointly models both the score ordinality and non-uniform intervals of proficiency labels. Extensive experiments on the TEEMI corpus show that our method consistently outperforms strong baselines and generalizes well to unseen prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。