arXiv:2512.21787cs.CL2025-12被引 2

针对方言阿拉伯语翻译难题,提出人性化后编辑评估框架

Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation

  • 构建五类错误分类体系与决策树标注协议
  • 发现方言术语与语义保留仍是主要挑战
  • 适合研究方言机器翻译与评估的学者使用

方言阿拉伯语到现代标准阿拉伯语(DA-MSA)翻译是机器翻译中的难点,因方言与标准语在词汇、句法和语义上差异显著。现有自动评估指标和通用人工评估框架难以捕捉方言特异性错误,制约了翻译评估进展。本文提出Ara-HOPE——一种以人为本的后编辑评估框架,包含五类错误分类体系和决策树标注协议。通过对比三个MT系统(阿拉伯语导向的Jais、通用GPT-3.5和基线NLLB-200),Ara-HOPE有效揭示了各系统间系统性性能差异。结果表明,方言特定术语与语义保留仍是DA-MSA翻译中最持久的挑战。Ara-HOPE为方言阿拉伯语翻译质量评估建立了新范式,并为提升方言感知型翻译系统提供可操作指导。相关标注文件与材料已公开于https://github.com/abdullahalabdullah/Ara-HOPE。

原文摘要 · Abstract (English)

Dialectal Arabic to Modern Standard Arabic (DA-MSA) translation is a challenging task in Machine Translation (MT) due to significant lexical, syntactic, and semantic divergences between Arabic dialects and MSA. Existing automatic evaluation metrics and general-purpose human evaluation frameworks struggle to capture dialect-specific MT errors, hindering progress in translation assessment. This paper introduces Ara-HOPE, a human-centric post-editing evaluation framework designed to systematically address these challenges. The framework includes a five-category error taxonomy and a decision-tree annotation protocol. Through comparative evaluation of three MT systems (Arabic-centric Jais, general-purpose GPT-3.5, and baseline NLLB-200), Ara-HOPE effectively highlights systematic performance differences between these systems. Our results show that dialect-specific terminology and semantic preservation remain the most persistent challenges in DA-MSA translation. Ara-HOPE establishes a new framework for evaluating Dialectal Arabic MT quality and provides actionable guidance for improving dialect-aware MT systems. For reproducibility, we make the annotation files and related materials publicly available at https://github.com/abdullahalabdullah/Ara-HOPE

机器翻译方言处理评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。