arXiv:2512.00938cs.CLcs.AI2025-12

针对阿拉伯语命名实体识别性能差的问题,提出可解释的组件化分析框架

DeformAr: Rethinking NER Evaluation through Component Analysis and Visual Analytics

  • 将语言系统拆解为数据与模型组件,分两阶段诊断性能差异
  • 通过交互式可视化发现模型在阿拉伯语中的错误模式与数据缺陷关联
  • 首个专为阿拉伯语设计的可解释性分析工具,适合低资源语言研究者

Transformer 模型在英语自然语言处理中表现优异,但在阿拉伯语命名实体识别(NER)任务上效果仍有限,即使使用更大规模预训练模型亦然。这一差距源于分词、数据质量及标注不一致等多重因素。现有研究多孤立分析各问题,未能揭示其联合影响。本文提出 DeformAr(基于 Transformer 的 NER 系统调试与评估框架),整合数据提取库与交互式仪表板,支持跨组件分析与行为分析两种模式。该框架将每种语言划分为数据与模型组件,分两阶段进行:第一阶段通过系统性诊断指标分析数据与模型子组件间的交互,回答‘是什么’‘如何’‘为何’;第二阶段结合可解释性技术、词粒度指标、交互式可视化与表示空间分析,实现组件感知的诊断流程,识别并解释模型行为与其底层表征模式和数据因素的关系。DeformAr 是首个面向阿拉伯语的组件化可解释性工具,为低资源语言模型分析提供关键支持。

原文摘要 · Abstract (English)

Transformer models have significantly advanced Natural Language Processing (NLP), demonstrating strong performance in English. However, their effectiveness in Arabic, particularly for Named Entity Recognition (NER), remains limited, even with larger pre-trained models. This performance gap stems from multiple factors, including tokenisation, dataset quality, and annotation inconsistencies. Existing studies often analyze these issues in isolation, failing to capture their joint effect on system behaviour and performance. We introduce DeformAr (Debugging and Evaluation Framework for Transformer-based NER Systems), a novel framework designed to investigate and explain the performance discrepancy between Arabic and English NER systems. DeformAr integrates a data extraction library and an interactive dashboard, supporting two modes of evaluation: cross-component analysis and behavioural analysis. The framework divides each language into dataset and model components to examine their interactions. The analysis proceeds in two stages. First, cross-component analysis provides systematic diagnostic measures across data and model subcomponents, addressing the "what," "how," and "why" behind observed discrepancies. The second stage applies behavioural analysis by combining interpretability techniques with token-level metrics, interactive visualisations, and representation space analysis. These stages enable a component-aware diagnostic process that detects model behaviours and explains them by linking them to underlying representational patterns and data factors. DeformAr is the first Arabic-specific, component-based interpretability tool, offering a crucial resource for advancing model analysis in under-resourced languages.

命名实体识别可解释性阿拉伯语组件分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。