arXiv:2503.19211cs.CL2025-03

构建阿拉伯语术语数据集,提升学术翻译一致性

MASRAD: Arabic Terminology Management Corpora with Semi-Automatic Construction

  • 通过自动提取阿拉伯语书籍中中外术语配对构建数据集
  • 最佳方法达到90.5%精确率和92.4%召回率
  • 适合语言工程与跨语言处理研究者使用

本文提出MASRAD,一个用于阿拉伯语术语管理的术语数据集及配套半自动构建方法。数据集条目为外文(非阿拉伯语)术语 $f$ 与其对应的阿拉伯语术语 $a$ 的配对,来源于专业、学术及领域专著。MASRAD-Ex作为构建第一步,自动提取阿拉伯语书籍中自然出现的术语对应关系,针对每个术语短语生成多个候选阿拉伯语词项,这些词项长度不一且紧邻外文术语。MASRAD-Ex计算词汇、语音、形态和语义相似度指标,采用启发式、机器学习及带后处理的机器学习方法筛选最优候选。经专家评审后,该数据集已公开。最佳方法实现90.5%精确率和92.4%召回率。

原文摘要 · Abstract (English)

This paper presents MASRAD, a terminology dataset for Arabic terminology management, and a method with supporting tools for its semi-automatic construction. The entries in MASRAD are $(f,a)$ pairs of foreign (non-Arabic) terms $f$, appearing in specialized, academic and field-specific books next to their Arabic $a$ counterparts. MASRAD-Ex systematically extracts these pairs as a first step to construct MASRAD. MASRAD helps improving term consistency in academic translations and specialized Arabic documents, and automating cross-lingual text processing. MASRAD-Ex leverages translated terms organically occurring in Arabic books, and considers several candidate pairs for each term phrase. The candidate Arabic terms occur next to the foreign terms, and vary in length. MASRAD-Ex computes lexicographic, phonetic, morphological, and semantic similarity metrics for each candidate pair, and uses heuristic, machine learning, and machine learning with post-processing approaches to decide on the best candidate. This paper presents MASRAD after thorough expert review and makes it available to the interested research community. The best performing MASRAD-Ex approach achieved 90.5% precision and 92.4% recall.

术语管理阿拉伯语数据集跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。