针对阿拉伯语方言翻译缺陷,提出六层语言学评估框架LQM。
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation

- 按语言学层级构建六级错误分类体系,覆盖从语用到拼写的全链路
- 基于3850句双语语料,标注6113个错误片段并生成带权重评分
- 适用于方言/文化敏感型翻译评估,可推广至其他多语种场景
现有机器翻译评估框架(包括自动指标与人类评估)普遍缺乏语言特异性,难以捕捉双语语言中因语言变体、内容覆盖和语用得体性不匹配导致的翻译失败。本文提出LQM:一种基于语言学的多维质量评估框架,通过六个语言学驱动的层级诊断翻译错误:社会语言学、语用学、语义学、形态句法、正字法与文字学。构建了一个包含3850句(每种方言550句)的双向平行语料库,涵盖七种阿拉伯语方言(埃及、阿联酋、约旦、毛里塔尼亚、摩洛哥、巴勒斯坦、也门),源自对话式且富含文化背景的内容。在零样本设置下评估六种大语言模型,并由专家进行逐跨度人工标注,共生成6113个标记错误段落,涉及3495个独特出错句子,同时提供严重度加权的质量评分。补充使用spBLEU作为自动指标。尽管验证于阿拉伯语,但该框架具有语言无关性,可轻松应用于或适配其他语言。LQM错误标注数据、提示模板与标注指南已公开于https://github.com/UBC-NLP/LQM_MT。
原文摘要 · Abstract (English)
Existing MT evaluation frameworks, including automatic metrics and human evaluation schemes such as Multidimensional Quality Metrics (MQM), are largely language-agnostic. However, they often fail to capture dialect- and culture-specific errors in diglossic languages (e.g., Arabic), where translation failures stem from mismatches in language variety, content coverage, and pragmatic appropriateness rather than surface form alone.We introduce LQM: Linguistically Motivated Multidimensional Quality Metrics for MT. LQM is a hierarchical error taxonomy for diagnosing MT errors through six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics (Figure 1). We construct a bidirectional parallel corpus of 3,850 sentences (550 per variety) spanning seven Arabic dialects (Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni), derived from conversational, culturally rich content. We evaluate six LLMs in a zero-shot setting and conduct expert span-level human annotation using LQM, producing 6,113 labeled error spans across 3,495 unique erroneous sentences, along with severity-weighted quality scores. We complement this analysis with an automatic metric (spBLEU). Though validated here on Arabic, LQM is a language-agnostic framework designed to be easily applied to or adapted for other languages. LQM annotated errors data, prompts, and annotation guidelines are publicly available at https://github.com/UBC-NLP/LQM_MT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。