小模型DIETA专注意英翻译,性能超多数同类模型。
DIETA: A Decoder-only transformer-based model for Italian-English machine TrAnslation
- 基于0.5亿参数的解码器仅架构,专为意英翻译设计。
- 在20700万句对上训练,测试中四项指标领先同规模模型。
- 公开数据集与评测集,适合做垂直领域机器翻译研究。
本文提出DIETA,一个拥有0.5亿参数的小型解码器仅Transformer模型,专门用于意大利语-英语机器翻译。我们构建并整理了一个涵盖议会记录、法律文本、网络爬取内容、字幕、新闻、文学等多领域的大型双语语料库,包含约20700万句对,并利用预训练模型生成3.52亿句回译数据。此外,我们创建并发布了新的小型评估集,由2025年WikiNews文章组成的450个句子,可用于评估当代文本的翻译质量。全面评估显示,DIETA在多个意英基准测试中表现优异,在32个系统的排行榜中始终位列第二四分位,且在五个测试套件中的四个上超越其他大多数参数少于30亿的模型。训练脚本、已训练模型、语料库及新评估集均公开可用,推动专用意英机器翻译的研究与发展。
原文摘要 · Abstract (English)
In this paper, we present DIETA, a small, decoder-only Transformer model with 0.5 billion parameters, specifically designed and trained for Italian-English machine translation. We collect and curate a large parallel corpus consisting of approximately 207 million Italian-English sentence pairs across diverse domains, including parliamentary proceedings, legal texts, web-crawled content, subtitles, news, literature and 352 million back-translated data using pretrained models. Additionally, we create and release a new small-scale evaluation set, consisting of 450 sentences, based on 2025 WikiNews articles, enabling assessment of translation quality on contemporary text. Comprehensive evaluations show that DIETA achieves competitive performance on multiple Italian-English benchmarks, consistently ranking in the second quartile of a 32-system leaderboard and outperforming most other sub-3B models on four out of five test suites. The training script, trained models, curated corpus, and newly introduced evaluation set are made publicly available, facilitating further research and development in specialized Italian-English machine translation. https://github.com/pkasela/DIETA-Machine-Translation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。