提出新训练方法,减少语音评分模型对作弊捷径的依赖。
Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers

- 设计新训练准则,抑制模型对输入特征的过度依赖
- 实验显示模型对可作弊特征的相关性显著降低
- 适合语言评估系统开发者与AI伦理研究者
越来越多的语音与语言处理任务直接使用音频或文本作为输入,而非提取特征。这类系统常采用复杂的变换器架构,能建立高度非线性的输入输出映射。然而,这些系统可能学习到‘捷径’——过度依赖输入中的特定部分来预测结果。在语言能力评估中,这种依赖使学习者可通过操纵特定特征提升分数,而非真正提高语言水平。本文提出一种新型训练准则,有效降低分类器对捷径的依赖,从而限制此类作弊行为。该方法在基于音频和语音识别文本的两种评估系统上验证,结果显示,改进前模型对可被利用特征的相关性高于人类参考值;引入新准则后,相关性降至接近参考水平。
原文摘要 · Abstract (English)
Increasingly, speech and language processing tasks take either audio or text directly rather than extracting features from these as the input to the classifier or regressor. Often these systems make use of complex, for example transformer-based, processes that have the ability to derive highly non-linear mappings between the input and the output. Unfortunately these systems can also learn ''shortcuts'' where the classifier is overly reliant on particular aspects of the input to yield the output. For the task of language proficiency assessment, this over-reliance can enable learners to increase their score by exploiting the shortcut rather than improving their ability. This paper introduces a novel training criterion that is able to reduce the classifier's reliance on shortcuts, thus for example limiting this option for malpractice in language assessment. This process is illustrated on two forms of assessment system, one based on the audio the other on the speech recognition text. The results show that, for both systems, there is higher correlations with features that could be exploited for malpractice than expected from the human reference, indicating an over-reliance on these features. By introducing the modified training criterion, this correlation can be reduced to be closer to the reference correlation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。