arXiv:2605.31452cs.CLcs.HC2026-05中稿 · EAMT-2026

为保密翻译场景测试本地大模型,发现部分表现可超专业系统。

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

论文配图:Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows
图 1 · 摘自论文原文
  • 用多语言语料在本地运行模型,不依赖云端
  • 最大本地模型在四方向上逼近甚至超过商用系统
  • 适合对隐私要求高的自由译员和小机构

基于前期工作,本文为自由译者及小型语言服务提供商开发了实用、低门槛的翻译技术评估方法。针对高度敏感的保密领域,提出离线翻译方案,避免使用云服务和商业大模型。将先前的三语语料库(RFTC)扩展为多语言语料库(RFMC),新增德语和简体中文对齐语料。在1000+句样本上,通过Ollama部署多个本地运行语言模型,对比四种语言方向表现。采用统一单提示调用,无微调或领域适配,评估对象包括商业NMT(DeepL、百度)、前沿大模型(GPT-5.2)及专业本地NMT系统(OPUS-CAT、NeuralDesktop、Promt)。自动评估使用MATEO。结果表明,本地大模型在不同语言方向与模型规模间表现差异显著。最佳本地模型在性能上达到或超越本地NMT系统,并接近前沿大模型,但仍落后于顶尖商业NMT系统。研究证实精心选择的本地大模型在隐私受限场景下具备可行性,为未来模型扩展与多语言能力研究提供依据。

原文摘要 · Abstract (English)

Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to evaluate translation technologies using rigorous yet accessible analytic methods. Here we address a high-stakes, specialized need: offline translation for confidentiality-sensitive domains in which privacy constraints preclude the use of cloud-based engines and commercial LLMs. We expand the Reeve Foundation Trilingual Corpus (RFTC) used in our previous work into a multilingual corpus (RFMC) by adding sentence-aligned German and Simplified Chinese reference translations. We then benchmark several locally runnable language models (via Ollama) across four language directions on 1000+ sentences selected from this corpus. We use consistent single-prompt calls without fine-tuning or domain adaptation, comparing local LLM outputs against commercial NMTs (DeepL, Baidu), a frontier LLM (GPT-5.2), and professional-grade local NMT systems (OPUS-CAT, NeuralDesktop, Promt). Automatic evaluation is conducted with MATEO. Results reveal substantial variation in local LLM performance across language directions and model sizes. The best local LLMs match or surpass local NMT systems and a frontier LLM, though they remain behind top commercial NMTs. These findings underscore the viability of carefully selected local LLM translation for privacy-constrained professionals and inform future research on model scaling and multilingual capability.

本地大模型翻译评估隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。