阿拉伯语NLP解释性研究存在方法、任务和语言三重短板,亟需更深入的本土化解释体系。
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
- 提出阿拉伯语XAI的五维分类框架,涵盖任务、方法、语言单位等
- 指出现有研究仅覆盖少数分类任务,生成与对话等任务解释不足
- 强调需针对阿拉伯语特有的语法、方言、文化等现象设计解释机制
可解释人工智能(XAI)已成为自然语言处理的重要议题,但阿拉伯语NLP在三个层面仍缺乏充分解释:一是方法差距,阿拉伯语XAI过度依赖LIME、SHAP、注意力可视化等少量事后解释技术,而主流NLP XAI已发展出诊断、反事实、探针、基于理由及以人为中心等多样化方法;二是任务差距,现有工作集中于情感分析、仇恨/攻击性语言检测、假新闻与垃圾信息分类,对生成、检索、翻译、摘要、结构化预测及对话等任务覆盖较弱;三是语言差距,多数解释聚焦影响词元,却很少解析阿拉伯语特有的形态、附着词、方言差异、文白异读、拼写歧义、符号标记、代码转换、命名实体、文化指涉以及古典与宗教语体等现象。本文通过系统综述文本、语音与多模态场景下的阿拉伯语XAI研究,主张阿拉伯语NLP不仅需要模型决策解释,更需忠实反映阿拉伯语作为语言、文化与社会技术对象的复杂性。我们构建了任务、方法、语言单位、语体、目标与评估实践的分类体系,并提出面向语言根基的阿拉伯语XAI研究议程。
原文摘要 · Abstract (English)
Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME, SHAP, attention visualization, and saliency, while broader NLP XAI offers richer diagnostic, counterfactual, probing, rationale-based, and human-centered methods. Second, there is a task gap: existing Arabic XAI work is concentrated in classification tasks, especially sentiment analysis, hate/offensive language detection, fake news, and spam, with weaker coverage of generation, retrieval, translation, summarization, structured prediction, and dialogue. Third, there is a linguistic gap: many explanations identify influential tokens, but rarely explain Arabic-specific phenomena such as morphology, clitics, dialectal variation, diglossia, orthographic ambiguity, diacritics, code-switching, named entities, cultural references, or Classical and religious registers. This critical structured survey synthesizes the reviewed literature on Arabic XAI across text, speech, and multimodal settings. We argue that Arabic NLP does not only need explanations of model decisions; it needs explanations that are faithful to Arabic as a linguistic, cultural, and sociotechnical object. We introduce a taxonomy of tasks, methods, linguistic units, varieties, goals, and evaluation practices, and propose a research agenda for linguistically grounded Arabic XAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。