构建了埃及阿拉伯语多文体双语数据集,助力机器翻译与语言研究。
ArzEn-MultiGenre: An aligned parallel dataset of Egyptian Arabic song lyrics, novels, and subtitles, with English translations
- 涵盖歌曲、小说、剧集字幕三类文本,人工对齐翻译。
- 含25,557个语段对,覆盖多种语言场景。
- 适合翻译研究、教学及专业译者参考使用。
ArzEn-MultiGenre 是一个包含埃及阿拉伯语歌曲歌词、小说和电视剧字幕的平行语料库,所有内容均经人工翻译并精确对齐英文版本。该数据集共包含25,557个语段对,可用于评估新型机器翻译模型、在少样本条件下微调大语言模型,以及优化谷歌翻译等商用翻译系统。此外,该数据集在翻译学、跨语言分析和词汇语义等领域具有重要研究价值,也可用于翻译教学训练及专业译者作为翻译记忆工具。其双重贡献在于:一是收录了现有埃及阿拉伯语-英语平行语料库中未涵盖的文本类型;二是由母语专家完成高质量人工翻译与对齐,具备金标准品质。
原文摘要 · Abstract (English)
ArzEn-MultiGenre is a parallel dataset of Egyptian Arabic song lyrics, novels, and TV show subtitles that are manually translated and aligned with their English counterparts. The dataset contains 25,557 segment pairs that can be used to benchmark new machine translation models, fine-tune large language models in few-shot settings, and adapt commercial machine translation applications such as Google Translate. Additionally, the dataset is a valuable resource for research in various disciplines, including translation studies, cross-linguistic analysis, and lexical semantics. The dataset can also serve pedagogical purposes by training translation students and aid professional translators as a translation memory. The contributions are twofold: first, the dataset features textual genres not found in existing parallel Egyptian Arabic and English datasets, and second, it is a gold-standard dataset that has been translated and aligned by human experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。