arXiv:2506.03593cs.CL2025-06ACL被引 1

对比语言学驱动与朴素数据增强,发现后者在多数情况下更有效。

Is linguistically-motivated data augmentation worth it?

  • 比较了语言学合理与朴素的数据增强方法在低资源语言上的效果。
  • 语言学驱动方法仅在生成样本不偏离原分布时才表现更好。
  • 适用于低资源语言的机器翻译与逐行注释任务研究者。

数据增强是应对数据稀缺的常用技术,通过生成合成数据来扩充训练集。已有研究表明,即使使用包含无意义词汇或违反语言规则的随机扰动数据,模型仍能获益。另一类方法生成符合语言规则的合成数据,需语言学知识且实现较复杂。然而,此前缺乏对两类方法系统的实证对比,无法判断语言学驱动增强是否真正提升下游性能。本文针对两种形态特征不同的低资源语言(Uspanteko 和 Arapaho),在机器翻译和逐行注释两项序列到序列任务上,系统评估多种增强策略及其组合。结果表明,语言学驱动方法虽有优势,但仅当生成样本与训练数据分布相近时才有效。

原文摘要 · Abstract (English)

Data augmentation, a widely-employed technique for addressing data scarcity, involves generating synthetic data examples which are then used to augment available training data. Researchers have seen surprising success from simple methods, such as random perturbations from natural examples, where models seem to benefit even from data with nonsense words, or data that doesn't conform to the rules of the language. A second line of research produces synthetic data that does in fact follow all linguistic constraints; these methods require some linguistic expertise and are generally more challenging to implement. No previous work has done a systematic, empirical comparison of both linguistically-naive and linguistically-motivated data augmentation strategies, leaving uncertainty about whether the additional time and effort of linguistically-motivated data augmentation work in fact yields better downstream performance. In this work, we conduct a careful and comprehensive comparison of augmentation strategies (both linguistically-naive and linguistically-motivated) for two low-resource languages with different morphological properties, Uspanteko and Arapaho. We evaluate the effectiveness of many different strategies and their combinations across two important sequence-to-sequence tasks for low-resource languages: machine translation and interlinear glossing. We find that linguistically-motivated strategies can have benefits over naive approaches, but only when the new examples they produce are not significantly unlike the training data distribution.

数据增强低资源语言机器翻译语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。