arXiv:2506.22038cs.CL2025-06中稿 · 2nd Workshop on Cr…被引 2

对比机器与人工翻译儿童文学,发现大模型更接近人类风格。

Can Peter Pan Survive MT? A Stylometric Study of LLMs, NMTs, and HTs in Children's Literature Translation

  • 构建彼得·潘中英儿童文学译本库,含人译、大模型和神经网络翻译各7版。
  • 大模型在创意文本特征上更接近人工译本,尤其在重复与节奏上表现更好。
  • 适合关注AI译文质量、语言风格迁移的研究者和教育从业者。

本研究从风格计量角度评估机器翻译(MT)与人工翻译(HT)在中英儿童文学翻译(CLT)中的表现。构建包含21个译本的彼得·潘语料库,涵盖7个人工翻译、7个大语言模型(LLM)翻译和7个神经机器翻译(NMT)输出。采用通用特征集(词汇、句法、可读性、n-gram等)与专用于创意文本翻译(CTT)的特征集(重复、韵律、可译性等),共提取447个语言学特征。通过机器学习分类与聚类分析发现:在通用特征上,人译与机译在连词分布及1词频-一扬比率上差异显著;NMT与LLM在描述性词语使用与副词比例上存在显著差异。在CTT特征上,LLM在分布上更接近人译,风格特征匹配度更高,显示出其在儿童文学翻译中的潜力。

原文摘要 · Abstract (English)

This study focuses on evaluating the performance of machine translations (MTs) compared to human translations (HTs) in English-to-Chinese children's literature translation (CLT) from a stylometric perspective. The research constructs a Peter Pan corpus, comprising 21 translations: 7 human translations (HTs), 7 large language model translations (LLMs), and 7 neural machine translation outputs (NMTs). The analysis employs a generic feature set (including lexical, syntactic, readability, and n-gram features) and a creative text translation (CTT-specific) feature set, which captures repetition, rhythm, translatability, and miscellaneous levels, yielding 447 linguistic features in total. Using classification and clustering techniques in machine learning, we conduct a stylometric analysis of these translations. Results reveal that in generic features, HTs and MTs exhibit significant differences in conjunction word distributions and the ratio of 1-word-gram-YiYang, while NMTs and LLMs show significant variation in descriptive words usage and adverb ratios. Regarding CTT-specific features, LLMs outperform NMTs in distribution, aligning more closely with HTs in stylistic characteristics, demonstrating the potential of LLMs in CLT.

儿童文学风格计量大模型翻译机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。