arXiv:2603.07372cs.CLcs.AI2026-03被引 1

针对低资源语种翻译,提出改进的模型适配方法提升质量评估准确性。

Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios

  • 在中间层使用低秩适配技术改进大模型,提升评估性能。
  • 在医疗等复杂领域,性能提升显著,最高达6.2%准确率增益。
  • 适用于低资源语言对和高风险应用场景的质量评估研究。

质量评估(QE)在无参考文本的场景中至关重要,尤其在领域特定和低资源语言场景下。本文研究了英译印地语四种领域(医疗、法律、旅游、通用)及五种语言对的句子级质量评估。系统比较了零样本、少样本与基于指南的提示策略,在选定的闭源与开源大模型上进行测试。结果表明,闭源模型仅通过提示即可实现强性能,而开源模型在高风险领域仍易受干扰。为此,我们采用基于低秩适配的ALOPE框架,并引入最新提出的低秩乘性适配(LoRMA)。实验显示,中间层适配显著提升评估效果,尤其在语义复杂的领域表现更优,为实际应用提供了更稳健的解决方案。代码与领域专用数据集已公开,以支持后续研究。

原文摘要 · Abstract (English)

Quality Estimation (QE) is essential for assessing machine translation quality in reference-less settings, particularly for domain-specific and low-resource language scenarios. In this paper, we investigate sentence-level QE for English to Indic machine translation across four domains (Healthcare, Legal, Tourism, and General) and five language pairs. We systematically compare zero-shot, few-shot, and guideline-anchored prompting across selected closed-weight and open-weight LLMs. Findings indicate that while closed-weight models achieve strong performance via prompting alone, prompt-only approaches remain fragile for open-weight models, especially in high-risk domains. To address this, we adopt ALOPE, a framework for LLM-based QE that uses Low-Rank Adaptation with regression heads attached to selected intermediate Transformer layers. We also extend ALOPE with recently proposed Low-Rank Multiplicative Adaptation (LoRMA). Our results show that intermediate-layer adaptation consistently improves QE performance, with gains in semantically complex domains, indicating a path toward more robust QE in practical scenarios. We release code and domain-specific QE datasets publicly to support further research.

机器翻译质量评估低资源大模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。