构建阿拉伯语跨目标立场检测基准数据集,助力模型泛化能力评估。
Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection

- 基于996条阿拉伯语推文,构建跨目标立场检测数据集
- 在女性驾车、电动车、三月制等目标上标注立场与情感
- 适合研究阿拉伯语自然语言处理与跨领域泛化的学者
公开可用的阿拉伯语目标特定立场检测数据集仍十分有限,尤其缺乏用于评估跨目标泛化的资源。本文提出 Mawqif-XT,包含从三个公开目标(女性驾车、电动车、三月制)收集的996条人工标注的阿拉伯语推文。每条推文均依据原始 Mawqif 标注方案标注立场、情感和讽刺标签。该扩展数据集作为独立测试集,用于评估模型对语义相关及未见目标的泛化能力,而原始 Mawqif 数据集则用于训练与开发。此外,我们使用多种阿拉伯语及多语言 Transformer 模型,以及零样本大语言模型(LLM)建立基线结果,以支持可复现评估。结合原始 Mawqif 数据集,Mawqif-v2 Extension 构成阿拉伯语立场检测中跨目标泛化的基准。
原文摘要 · Abstract (English)
Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-XT, consisting of 996 manually annotated Arabic tweets collected from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is annotated with stance, sentiment, and sarcasm labels following the original Mawqif annotation scheme. The released extension is intended as a held-out evaluation set for assessing model generalization to both semantically related and previously unseen targets, while the original Mawqif dataset is used for training and development. In addition, we establish baseline results using several Arabic and multilingual transformer models, as well as zero-shot large language models (LLMs), to facilitate reproducible evaluation. Together with the original Mawqif dataset, the Mawqif-v2 Extension provides a benchmark for evaluating cross-target generalization in Arabic stance detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。