用多大模型投票蒸馏,让小模型实现高效精准翻译质量评估
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation

- 用多个大模型组成评审团生成高质量合成数据
- 小模型经蒸馏后性能远超普通小模型,接近大模型水平
- 适合需要快速、低成本部署翻译质检系统的团队
大型语言模型(LLMs)在基于MQM的翻译质量(TQ)评估中表现优异,而大型推理模型(LRMs)的进展更带来提升潜力。然而,两者计算成本高,难以规模化部署;小型语言模型(SLMs)虽高效,却难以应对评估任务所需的复杂推理。本文通过广泛实证研究,在多种TQ评估设置下对比了SLMs、LLMs和LRMs,全面呈现当前技术格局并确立最佳实践。为解决可扩展性问题,我们提出TQLite——一种新型蒸馏框架,使SLMs能逼近最优LRM评估器的性能。该方法利用多LRM评审团,通过实用的数据清洗与模型响应聚合技术生成高质量合成训练数据。实验表明,经TQLite训练的SLMs在MQM评估上表现强劲,显著优于标准SLMs,为LLM和LRM评估器提供了可扩展、低成本的替代方案。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. However, both LLMs and LRMs are computationally expensive to deploy at scale, while small language models (SLMs)---though much more efficient---struggle with the complex reasoning required for evaluation tasks. In this work, we present an extensive empirical study benchmarking SLMs, LLMs, and LRMs across a wide range of TQ evaluation setups, providing a comprehensive view of the current landscape and establishing best practices. To address the scalability challenge, we introduce TQLite, a novel distillation framework that enables SLMs to approach the MQM evaluation performance of the best LRM-based evaluators. Our approach leverages a multi-LRM jury to generate high-quality synthetic training data via practical data curation techniques and aggregation of evaluation responses across a diverse panel of models. Our results demonstrate that SLMs trained via TQLite achieve strong MQM evaluation performance that far exceeds off-the-shelf evaluation capabilities of standard SLMs, offering a scalable and cost-effective alternative to LLM- and LRM-based evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。