研究合成数据质量对多语言切换任务的影响,发现效果因任务而异。
The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR
- 对比多种数据增强方法在多语言切换中的表现
- 不同任务下合成数据质量提升效果差异显著
- 适合关注多语言技术落地的研究者参考
代码切换(code-switching)作为全球普遍现象,亟需在语言技术中加以应对。当前主要瓶颈是数据稀缺,促使研究转向代码切换数据增强。然而现有文献缺乏系统性研究来揭示合成数据质量与NLP任务性能之间的关系。本文在机器翻译(MT)基础上,扩展至自动语音识别(ASR)与级联语音翻译(ST),测试结论的通用性。实验涵盖词汇替换、语言学理论和回译等多种增强技术。基于MT、ASR与ST的结果,得出关于不同增强方法有效性及数据质量影响的洞见。
原文摘要 · Abstract (English)
Code-switching, the act of alternating between languages, emerged as a prevalent global phenomenon that needs to be addressed for building user-friendly language technologies. A main bottleneck in this pursuit is data scarcity, motivating research in the direction of code-switched data augmentation. However, current literature lacks comprehensive studies that enable us to understand the relation between the quality of synthetic data and improvements on NLP tasks. We extend previous research conducted in this direction on machine translation (MT) with results on automatic speech recognition (ASR) and cascaded speech translation (ST) to test generalizability of findings. Our experiments involve a wide range of augmentation techniques, covering lexical replacements, linguistic theories, and back-translation. Based on the results of MT, ASR, and ST, we draw conclusions and insights regarding the efficacy of various augmentation techniques and the impact of quality on performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。