自监督语音模型需任务微调才能识别汉语声调语境补偿
Perceptual compensation for tonal context in self-supervised speech models

- 用伪重复实验测试wav2vec2.0对汉语声调语境的补偿能力
- 纯预训练模型无补偿表现,微调后分类器略有提升但未达人类水平
- 提示:语音建模中监督目标对音位规律抽象至关重要
本研究考察了wav2vec2.0架构在多音调语言(汉语普通话)中是否具备对音位上下文的感知补偿能力。通过复现感知补偿实验,比较了纯自监督预训练模型与针对普通话语音识别微调后的模型在嵌入相似性及探针分类器输出上的表现。结果表明,纯预训练模型在嵌入相似性上未体现补偿效应;探针分类器虽显示出一定程度的补偿迹象,并随层数加深分类能力提升,但在孤立音节测试中未能复现人类表现。研究结果与此前关于预训练可自发产生音位结构敏感性的报告相悖,暗示监督目标对于促进至少部分音位规律抽象是必要的。
原文摘要 · Abstract (English)
This study examines the extent to which the wav2vec2.0 architecture exhibits evidence of compensation for phonological context. We conducted a pseudo-replication of a perceptional compensation experiment on Mandarin Chinese tones, and compared the embedding similarities and probing classifier outputs between a purely self-supervised pre-trained model and a model fine-tuned for Mandarin ASR. No evidence of compensation was found in the embedding similarities of the purely pre-trained model. Probing classifiers showed some evidence of compensation in addition to the expected layer-wise improvements in categorization, but failed to replicate human performance on isolated test syllables. Our findings contrast with previous reports of sensitivity to phonological structure emerging through pre-training alone, and suggest that supervised objectives may be necessary to encourage the abstraction of at least some types of phonological regularities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。