多任务学习会损害二语语音识别的发音转写效果,因表征纠缠所致。
Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

- 通过分析编码器表征纠缠,揭示多任务学习在韩英双语中的差异表现
- 英语发音转写准确率随表面与语义差异增大而显著下降
- 适合关注多任务学习缺陷及二语语音识别优化的研究者
二语(L2)语音识别通常需要同时输出发音转写和意图语义。多任务学习(MTL)是自然选择,因其假设共享表征对两个输出均有帮助。然而本文发现,该假设在韩语与英语之间不成立:MTL虽提升语义识别性能,却严重降低发音转写准确率,尤其在英语中,其下降程度与莱文斯坦编辑距离衡量的表面-语义差异呈正相关。编码器分析表明,这种现象源于编码器层级的表征纠缠——韩语保持解耦表征,而英语生成几乎相同的表征。跨输出解码器分析显示,语义双输出解码器能适应独特表征,而发音双输出解码器仍受编码器约束。这些发现提示应设计减轻编码器层级纠缠的MTL框架,以缓解双输出二语语音识别中的发音退化问题。
原文摘要 · Abstract (English)
Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural approach because it assumes that shared representations benefit both outputs. However, this paper shows that this assumption does not hold across Korean and English. MTL improves meaning but degrades surface transcription, especially in English, where the degradation scales with surface-meaning divergence measured by Levenshtein edit distance. Encoder analysis links these patterns to encoder-level entanglement, with Korean preserving disentangled representations while English produces nearly identical ones. Cross-output decoder analysis shows that the meaning dual-output decoder adapts with a unique representation, while the surface dual-output decoder remains constrained by the encoder. These findings motivate the design of MTL frameworks that mitigate encoder-level entanglement to reduce surface degradation in dual-output L2 automatic speech recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。