arXiv:2411.13775cs.CLcs.AI2024-11被引 12

GPT-4翻译水平相当于初级译员,跨语言表现稳定但有刻板和不一致问题。

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

  • 对比不同水平人类译者,评估GPT-4在多语言多领域表现。
  • GPT-4总错误数与初级译员相当,但低于资深译员。
  • 适合关注大模型翻译能力边界的研究者与从业者。

本研究系统评估了GPT-4与不同专业水平人类译者的翻译能力。通过基于MQM标准的人工评估,考察了中英、俄英、中文-印地语三组语言对,以及新闻、科技、生物医学三个领域。结果表明,GPT-4的总错误数与初级译员相当,仍不及资深译员;与传统神经机器翻译不同,其在资源匮乏语言方向未出现明显性能下降。定性分析显示,GPT-4倾向过度直译且词汇不一致,而人类译者有时会过度解读上下文并引入幻觉。这是首次在不同专业水平下系统比较大语言模型与人类译者的综合性研究,揭示了当前大模型翻译系统的实际能力与局限。

原文摘要 · Abstract (English)

This study presents a comprehensive evaluation of GPT-4's translation capabilities compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translations across three language pairs (Chinese$\longleftrightarrow$English, Russian$\longleftrightarrow$English, and Chinese$\longleftrightarrow$Hindi) and three domains (News, Technology, and Biomedical). Our findings reveal that GPT-4 achieves performance comparable to junior-level translators in terms of total errors, while still lagging behind senior translators. Unlike traditional Neural Machine Translation systems, which show significant performance degradation in resource-poor language directions, GPT-4 maintains consistent translation quality across all evaluated language pairs. Through qualitative analysis, we identify distinctive patterns in translation approaches: GPT-4 tends toward overly literal translations and exhibits lexical inconsistency, while human translators sometimes over-interpret context and introduce hallucinations. This study represents the first systematic comparison between LLM and human translators across different proficiency levels, providing valuable insights into the current capabilities and limitations of LLM-based translation systems.

大模型翻译人类对比多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。