小模型也能精准评估翻译质量,还支持错误定位与修改建议。
CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs

- 用小于300亿参数的开源小模型实现单次提示评估
- 相关性超越传统指标和人工标注者一致率
- 适合关注隐私与成本的翻译质检场景
当前机器翻译的质量评估(QE)依赖大规模专有大模型,引发数据隐私担忧。我们证明,参数量小于300亿的开源小模型是可行、低成本且保护隐私的替代方案。通过单次提示策略,模型可同时生成质量评分、MQM错误标注、建议修正及完整修订文本。分析显示,这些模型在系统级相关性上达到与人类判断高度匹配的水平,优于传统神经指标、微调模型以及人工标注者间的一致性,有效逼近大型专有大模型的能力。
原文摘要 · Abstract (English)
Current state-of-the-art Quality Estimation (QE) in machine translation relies on massive, proprietary LLMs, raising data privacy concerns. We demonstrate that smaller, open-source LLMs (<30B parameters) are a viable, cost-effective and privacy-preserving alternative. Using a single-pass prompting strategy, our models simultaneously generate quality scores, MQM error annotations, suggested error corrections, and full post-editions. Our analysis shows these models achieve highly competitive system-level correlations with human judgments that outperform traditional neural metrics, fine-tuned models, and human inter-annotator agreement, effectively approximating the capabilities of much larger proprietary LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。