arXiv:2505.01560cs.CL2025-05

大模型翻译虽更准确,但成本高出数倍,未必值得当前投入。

AI agents may be worth the hype but not the resources (yet): An initial exploration of machine translation quality and costs in three language pairs in the legal and news domains

  • 对比五种翻译方法,包括单次调用和多智能体流程。
  • 大模型输出更自然流畅,但耗能是传统翻译的5至15倍。
  • 适合关注质量与成本平衡的研究者或实际部署场景。

大型语言模型(LLMs)与多智能体协同被视为机器翻译的新突破,但其相对于传统神经机器翻译(NMT)的优势尚不明确。本文进行了实证检验:在法律合同与新闻文本三组英源语言对(西语、加泰罗尼亚语、土耳其语)上,对比了五种范式——Google Translate(强基准NMT)、GPT-4o(通用大模型)、o1-preview(推理增强大模型)及两种基于GPT-4o的智能体工作流(串行三阶段与迭代优化)。自动评估采用COMET、BLEU、chrF2、TER;人工评估由专家评分语义恰当性与流畅度;效率以输入输出总token数映射2025年4月定价。自动指标显示成熟NMT系统在十二项组合中排名第一(七项),o1-preview在其余多数情况并列或位列第二,而多智能体流程落后。人工评估则部分逆转此趋势:o1-preview在六项比较中五项表现最优,迭代智能体一次领先,表明推理层捕捉到表面指标未体现的语义细微差别。然而这些质量提升代价高昂:串行智能体消耗约五倍、迭代智能体达十五倍于NMT或单次大模型的token量。论文倡导多维度、成本敏感的评估体系,并提出轻量化协调、选择性激活及混合流水线等研究方向,或可打破当前权衡格局。

原文摘要 · Abstract (English)

Large language models (LLMs) and multi-agent orchestration are touted as the next leap in machine translation (MT), but their benefits relative to conventional neural MT (NMT) remain unclear. This paper offers an empirical reality check. We benchmark five paradigms, Google Translate (strong NMT baseline), GPT-4o (general-purpose LLM), o1-preview (reasoning-enhanced LLM), and two GPT-4o-powered agentic workflows (sequential three-stage and iterative refinement), on test data drawn from a legal contract and news prose in three English-source pairs: Spanish, Catalan and Turkish. Automatic evaluation is performed with COMET, BLEU, chrF2 and TER; human evaluation is conducted with expert ratings of adequacy and fluency; efficiency with total input-plus-output token counts mapped to April 2025 pricing. Automatic scores still favour the mature NMT system, which ranks first in seven of twelve metric-language combinations; o1-preview ties or places second in most remaining cases, while both multi-agent workflows trail. Human evaluation reverses part of this narrative: o1-preview produces the most adequate and fluent output in five of six comparisons, and the iterative agent edges ahead once, indicating that reasoning layers capture semantic nuance undervalued by surface metrics. Yet these qualitative gains carry steep costs. The sequential agent consumes roughly five times, and the iterative agent fifteen times, the tokens used by NMT or single-pass LLMs. We advocate multidimensional, cost-aware evaluation protocols and highlight research directions that could tip the balance: leaner coordination strategies, selective agent activation, and hybrid pipelines combining single-pass LLMs with targeted agent intervention.

机器翻译大模型成本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。