用Transformer提升英印法律文本翻译,微调效果优于从零训练。
From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation
- 微调预训练的OPUS-MT模型,实现法律领域适配。
- 微调模型达46.03分的SacreBLEU,显著优于基线和从零训练。
- 适合关注法律翻译与多语言司法公平的研究者。
在印度等多语言国家,法律信息常因语言障碍难以获取,大量司法文档仍以英语呈现。法律机器翻译(L-MT)可提供可扩展的解决方案,实现法律文件的准确、可访问翻译。本文参与JUST-NLP 2025法律机器翻译共享任务,聚焦英语-印地语翻译,采用基于Transformer的两种互补策略:对预训练的OPUS-MT模型进行领域微调,以及使用提供的法律语料库从零训练Transformer模型。性能通过SacreBLEU、chrF++、TER、ROUGE、BERTScore、METEOR和COMET等标准指标评估。微调后的OPUS-MT模型取得46.03的SacreBLEU得分,显著优于基线模型和从零训练模型。结果表明,领域适配能有效提升翻译质量,展示了L-MT系统在改善多语言环境中司法可及性与透明度方面的潜力。
原文摘要 · Abstract (English)
In multilingual nations like India, access to legal information is often hindered by language barriers, as much of the legal and judicial documentation remains in English. Legal Machine Translation (L-MT) offers a scalable solution to this challenge by enabling accurate and accessible translations of legal documents. This paper presents our work for the JUST-NLP 2025 Legal MT shared task, focusing on English-Hindi translation using Transformer-based approaches. We experiment with 2 complementary strategies, fine-tuning a pre-trained OPUS-MT model for domain-specific adaptation and training a Transformer model from scratch using the provided legal corpus. Performance is evaluated using standard MT metrics, including SacreBLEU, chrF++, TER, ROUGE, BERTScore, METEOR, and COMET. Our fine-tuned OPUS-MT model achieves a SacreBLEU score of 46.03, significantly outperforming both baseline and from-scratch models. The results highlight the effectiveness of domain adaptation in enhancing translation quality and demonstrate the potential of L-MT systems to improve access to justice and legal transparency in multilingual contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。