为电商等工业场景设计了首个公开翻译评测基准,解决通用模型在专业领域表现不佳的问题。
TransBench: Benchmarking Machine Translation for Industrial-Scale Applications
- 构建三层次评估框架:语言基础、领域专精与文化适配
- 发布含1.7万句的电商翻译数据集,覆盖33种语言对和4类场景
- 引入新指标Marco-MOS,支持可复现的工业级翻译质量评估
机器翻译已成为全球化产业中跨边界沟通的关键工具,尤其在电子商务、金融和法律服务等领域。尽管大语言模型显著提升了翻译质量,但在工业应用场景中,通用模型因缺乏领域术语、文化差异和风格规范而表现受限。现有评估框架难以有效衡量特定情境下的性能,导致学术评测与实际效果脱节。为此,我们提出一个三层翻译能力框架:(1) 基础语言能力,(2) 领域专精能力,(3) 文化适应能力,强调多维度综合评估。我们推出TransBench,一个面向工业级翻译的基准,初期聚焦国际电子商务,包含17,000条专业译文,覆盖4个主要场景和33种语言对。该基准融合传统指标(BLEU、TER)与新型领域专用评估模型Marco-MOS,提供可复现的基准构建指南。贡献包括:(1) 工业翻译评估的结构化框架,(2) 首个公开可用的电商翻译基准,(3) 用于多层级质量检测的新指标,(4) 开源评估工具。本工作弥合了评估差距,助力研究者与从业者系统性提升工业场景下的MT系统性能。
原文摘要 · Abstract (English)
Machine translation (MT) has become indispensable for cross-border communication in globalized industries like e-commerce, finance, and legal services, with recent advancements in large language models (LLMs) significantly enhancing translation quality. However, applying general-purpose MT models to industrial scenarios reveals critical limitations due to domain-specific terminology, cultural nuances, and stylistic conventions absent in generic benchmarks. Existing evaluation frameworks inadequately assess performance in specialized contexts, creating a gap between academic benchmarks and real-world efficacy. To address this, we propose a three-level translation capability framework: (1) Basic Linguistic Competence, (2) Domain-Specific Proficiency, and (3) Cultural Adaptation, emphasizing the need for holistic evaluation across these dimensions. We introduce TransBench, a benchmark tailored for industrial MT, initially targeting international e-commerce with 17,000 professionally translated sentences spanning 4 main scenarios and 33 language pairs. TransBench integrates traditional metrics (BLEU, TER) with Marco-MOS, a domain-specific evaluation model, and provides guidelines for reproducible benchmark construction. Our contributions include: (1) a structured framework for industrial MT evaluation, (2) the first publicly available benchmark for e-commerce translation, (3) novel metrics probing multi-level translation quality, and (4) open-sourced evaluation tools. This work bridges the evaluation gap, enabling researchers and practitioners to systematically assess and enhance MT systems for industry-specific needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。