构建首个大规模阿拉伯语礼貌性数据集,助力文化敏感型自然语言处理研究。
ADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational Sociopragmatics
- 基于阿拉伯语言传统与语用理论,标注多方言文本的礼貌、不礼貌、中性三类标签。
- 包含1万条样本,覆盖4种方言和16类礼貌特征,标注一致性kappa达0.703。
- 支持从传统模型到大模型的全面评测,适合研究跨文化语用学的学者使用。
随着对文化敏感型自然语言处理系统的需求日益增长,跨语言社会语用现象资源的重要性凸显。然而,阿拉伯语礼貌性检测资源仍严重不足,尽管阿拉伯语交流中蕴含丰富的礼貌表达。本文提出ADAB(阿拉伯礼貌性数据集),从社交媒体、电商及客服等四个在线平台收集数据,涵盖现代标准阿拉伯语及海湾、埃及、黎凡特、马格里布四种方言。数据基于阿拉伯语言传统与语用理论进行标注,分为礼貌、不礼貌、中性三类,共10,000个样本,包含16类礼貌特征的语义标注,标注者间一致性kappa=0.703。我们对40种模型配置(包括传统机器学习、Transformer模型及大语言模型)进行了基准测试,旨在推动面向礼貌性的阿拉伯语NLP研究。
原文摘要 · Abstract (English)
The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for politeness detection remain under-explored, despite the rich and complex politeness expressions embedded in Arabic communication. In this paper, we introduce ADAB (Arabic Politeness Dataset), a new annotated Arabic dataset collected from four online platforms, including social media, e-commerce, and customer service domains, covering Modern Standard Arabic and multiple dialects (Gulf, Egyptian, Levantine, and Maghrebi). The dataset was annotated based on Arabic linguistic traditions and pragmatic theory, resulting in three classes: polite, impolite, and neutral. It contains 10,000 samples with linguistic feature annotations across 16 politeness categories and achieves substantial inter-annotator agreement (kappa = 0.703). We benchmark 40 model configurations, including traditional machine learning, transformer-based models, and large language models. The dataset aims to support research on politeness-aware Arabic NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。