arXiv:2409.09568cs.CL2024-09Conference of the …

警惕NLP工具削弱文本多样性,影响语言表达丰富性

Thesis proposal: Are We Losing Textual Diversity to Natural Language Processing?

  • 用统计特征量化文本多样性,分析NMT系统对异常文本的处理偏差
  • 实验发现机器翻译比人工翻译更倾向简化语言结构,降低多样性
  • 适合关注AI语言生成质量与社会影响的研究者和从业者

本文探讨当前广泛使用的自然语言处理算法在处理和生成文本时可能存在的局限性。以神经机器翻译(NMT)为研究场景,提出这些算法可能因固有归纳偏置,在处理非典型文本时产生负面影响。通过定义多尺度(句子、语篇、语言层)的文本多样性度量方法,基于词级意外性分布的均匀性或节奏性进行量化。实验表明,相比人类译者,NMT系统在翻译过程中会显著降低文本多样性,导致输出语言趋于同质化。研究进一步探究训练目标与解码策略是造成该现象的潜在原因。最终目标是开发新型模型,避免强制输出中统计特性的均匀分布,支持更具全局规划能力的翻译,以应对翻译任务本身的内在模糊性。

原文摘要 · Abstract (English)

This thesis argues that the currently widely used Natural Language Processing algorithms possibly have various limitations related to the properties of the texts they handle and produce. With the wide adoption of these tools in rapid progress, we must ask what these limitations are and what are the possible implications of integrating such tools even more deeply into our daily lives. As a testbed, we have chosen the task of Neural Machine Translation (NMT). Nevertheless, we aim for general insights and outcomes, applicable even to current Large Language Models (LLMs). We ask whether the algorithms used in NMT have inherent inductive biases that are beneficial for most types of inputs but might harm the processing of untypical texts. To explore this hypothesis, we define a set of measures to quantify text diversity based on its statistical properties, like uniformity or rhythmicity of word-level surprisal, on multiple scales (sentence, discourse, language). We then conduct a series of experiments to investigate whether NMT systems struggle with maintaining the diversity of such texts, potentially reducing the richness of the language generated by these systems, compared to human translators. We search for potential causes of these limitations rooted in training objectives and decoding algorithms. Our ultimate goal is to develop alternatives that do not enforce uniformity in the distribution of statistical properties in the output and that allow for better global planning of the translation, taking into account the intrinsic ambiguity of the translation task.

NLP文本多样性机器翻译语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。