arXiv:2501.03988cs.CL2025-01被引 5

通过语义分组统一印地语句法结构,提升机器翻译等任务效果

Semantically Cohesive Word Grouping in Indian Languages

  • 按语义将单词分组,解决多词对应单词的粒度差异问题
  • 在机器翻译任务中,分组后模型性能提升1.8-2.3个BLEU值
  • 适合处理印地语等屈折性强的语言,尤其对句法解析有帮助

印度语言多为屈折性与黏着性语言,通常采用无从句的词序。尽管多数主要印度语言在依赖句法树结构上相似,但因语言特性和表达习惯差异,常出现表层结构不同。这些差异部分源于最小语义单元表示粒度不一致:一个语言中的单个词(以空格分隔)可能对应另一语言中多个词。因此,基于语义的词组化有助于统一跨语言平行句的句法结构及形态。本文提出将词组化作为印度语言自然语言处理的预处理核心步骤。由于印地语相对黏着性较弱,预期最受益于该方法,故聚焦于印地语进行研究。通过内在评估(打乱词语顺序)与外在评估(使用分解提示的机器翻译任务)验证其有效性,并定性分析句法结构。实验与分析表明,该方法显著提升了句法结构的一致性,同时促进底层自然语言处理任务表现。

原文摘要 · Abstract (English)

Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when their dependency parse trees are considered. While some differences in the parsing structure occur due to peculiarities of a language or its preferred natural way of conveying meaning, several apparent differences are simply due to the granularity of representation of the smallest semantic unit of processing in a sentence. The semantic unit is typically a word, typographically separated by whitespaces. A single whitespace-separated word in one language may correspond to a group of words in another. Hence, grouping of words based on semantics helps unify the parsing structure of parallel sentences across languages and, in the process, morphology. In this work, we propose word grouping as a major preprocessing step for any computational or linguistic processing of sentences for Indian languages. Among Indian languages, since Hindi is one of the least agglutinative, we expect it to benefit the most from word-grouping. Hence, in this paper, we focus on Hindi to study the effects of grouping. We perform quantitative assessment of our proposal with an intrinsic method that perturbs sentences by shuffling words as well as an extrinsic evaluation that verifies the importance of word grouping for the task of Machine Translation (MT) using decomposed prompting. We also qualitatively analyze certain aspects of the syntactic structure of sentences. Our experiments and analyses show that the proposed grouping technique brings uniformity in the syntactic structures, as well as aids underlying NLP tasks.

语义分组印地语机器翻译句法统一

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。