首个约鲁巴语讽刺检测数据集,解决非洲低资源语言语义理解难题
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
- 构建约鲁巴语讽刺数据集,结合文化背景设计标注协议
- 三位母语者达成近乎完美一致(κ=0.766),83.3%样本完全共识
- 适合研究非洲语言语义、跨文化NLP及低资源语言模型
讽刺检测在计算语义中面临核心挑战,需识别字面意义与实际意图之间的差异。该挑战在低资源语言中尤为突出,因缺乏标注数据。本文提出首个约鲁巴语讽刺检测的黄金标准数据集Yor-Sarc,约鲁巴语是使用人数超过5000万的声调尼日尔-刚果语系语言。数据集包含436个实例,由三位来自不同方言背景的母语者根据专为约鲁巴语讽刺设计的标注协议进行标注,该协议融合上下文敏感解读与社区导向指南,并附有详尽的标注者间一致性分析,以支持其他非洲语言的复现。标注一致性达到显著至几乎完美水平(Fleiss' κ=0.7660;成对Cohen's κ=0.6732–0.8743),83.3%的样本达成完全一致。其中一对标注者达成近乎完美一致性(κ=0.8743;原始一致率93.8%),超越多项英语讽刺研究基准。剩余16.7%多数同意案例保留为软标签,用于不确定性感知建模。Yor-Sarc(https://github.com/toheebadura/yor-sarc)有望推动低资源非洲语言的语义解析与文化敏感型自然语言处理研究。
原文摘要 · Abstract (English)
Sarcasm detection poses a fundamental challenge in computational semantics, requiring models to resolve disparities between literal and intended meaning. The challenge is amplified in low-resource languages where annotated datasets are scarce or nonexistent. We present \textbf{Yor-Sarc}, the first gold-standard dataset for sarcasm detection in Yorùbá, a tonal Niger-Congo language spoken by over $50$ million people. The dataset comprises 436 instances annotated by three native speakers from diverse dialectal backgrounds using an annotation protocol specifically designed for Yorùbá sarcasm by taking culture into account. This protocol incorporates context-sensitive interpretation and community-informed guidelines and is accompanied by a comprehensive analysis of inter-annotator agreement to support replication in other African languages. Substantial to almost perfect agreement was achieved (Fleiss' $κ= 0.7660$; pairwise Cohen's $κ= 0.6732$--$0.8743$), with $83.3\%$ unanimous consensus. One annotator pair achieved almost perfect agreement ($κ= 0.8743$; $93.8\%$ raw agreement), exceeding a number of reported benchmarks for English sarcasm research works. The remaining $16.7\%$ majority-agreement cases are preserved as soft labels for uncertainty-aware modelling. Yor-Sarc\footnote{https://github.com/toheebadura/yor-sarc} is expected to facilitate research on semantic interpretation and culturally informed NLP for low-resource African languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。