arXiv:2604.00015cs.CL2026-04

构建首个面向科学文本的阿拉伯语-英语双语评估语料库,提升机器翻译评测精度。

ASCAT: An Arabic Scientific Corpus and Benchmark for Advanced Translation Evaluation

  • 通过多引擎翻译+专家校验,构建涵盖5个学科的长篇科学摘要语料
  • 包含67,293个英文词和60,026个阿拉伯词语,平均每篇141.7词(英)/111.78词(阿)
  • 支持大模型科学翻译质量评估,适合研究阿拉伯语机器翻译的学者使用

我们提出ASCAT(阿拉伯语科学语料库用于高级翻译评估),一个高质量的英阿平行语料库,专为科学翻译评估而设计。该语料库通过系统化多引擎翻译与人工验证流程构建,不同于现有依赖短句或单一领域文本的阿拉伯语-英语语料库,ASCAT聚焦于完整科学摘要,平均长度为141.7词(英文)和111.78词(阿拉伯文),覆盖物理、数学、计算机科学、量子力学及人工智能五个领域。每篇摘要采用三种互补架构生成:生成式AI(Gemini)、基于Transformer的模型(Hugging Face quickmt-en-ar)以及商用MT API(Google Translate、DeepL),随后由领域专家在词汇、语法和语义层面进行验证。最终语料库包含67,293个英文词元和60,026个阿拉伯词语元,阿拉伯语词汇量达17,604个独特词,反映语言形态丰富性。我们在该语料库上对三个先进大模型进行基准测试:GPT-4o-mini(BLEU: 37.07)、Gemini-3.0-Flash-Preview(BLEU: 30.44)、Qwen3-235B-A22B(BLEU: 23.68),验证其判别力。ASCAT填补了阿拉伯语科学机器翻译资源的空白,旨在支持科学翻译质量的严谨评估与特定领域模型的训练。

原文摘要 · Abstract (English)

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine translation and human validation pipeline. Unlike existing Arabic-English corpora that rely on short sentences or single-domain text, ASCAT targets full scientific abstracts averaging 141.7 words (English) and 111.78 words (Arabic), drawn from five scientific domains: physics, mathematics, computer science, quantum mechanics, and artificial intelligence. Each abstract was translated using three complementary architectures generative AI (Gemini), transformer-based models (Hugging Face \texttt{quickmt-en-ar}), and commercial MT APIs (Google Translate, DeepL) and subsequently validated by domain experts at the lexical, syntactic, and semantic levels. The resulting corpus contains 67,293 English tokens and 60,026 Arabic tokens, with an Arabic vocabulary of 17,604 unique words reflecting the morphological richness of the language. We benchmark three state-of-the-art LLMs on the corpus GPT-4o-mini (BLEU: 37.07), Gemini-3.0-Flash-Preview (BLEU: 30.44), and Qwen3-235B-A22B (BLEU: 23.68) demonstrating its discriminative power as an evaluation benchmark. ASCAT addresses a critical gap in scientific MT resources for Arabic and is designed to support rigorous evaluation of scientific translation quality and training of domain-specific translation models.

机器翻译阿拉伯语科学文本评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。