发现标注时间越接近,情感标注一致性越高,可用来提升数据质量。
Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

- 用时间同步性预测标注一致性,发现1分钟内标注κ=0.98
- 跨批次标注一致性下降超32点,主要在负面与中性边界出错
- 适合关注非洲语言NLP数据质量的科研人员使用
当标注任务持续数周且标注员较少时,保持高质量标注极为困难。本文构建了一个包含3,565条推文的塞茨瓦纳语情感数据集,由三位母语标注员分八批完成标注。尽管总体随机一致性系数κ=0.76(优秀),但各批次κ值下降超过32点。通过六项分析发现:(i) 标注混淆集中于负面/中性边界;(ii) 两名标注员出现符合自动化标注特征的连续标注偏差;(iii) 时间同步性是κ的主导预测因子——同一分钟内标注的κ达0.98,而相隔超过一天的仅0.65。标注速度和文本特征对κ无显著影响。在三分类情感分类上,微调三种开源多语言编码器及专有模型(GPT-5、Gemini),F1得分提升29至43点,其中GPT-5少样本表现最佳(62.2宏F1)。数据集、每条标注的时间戳及分析代码已公开,支持未来非洲语言NLP资源的质量可复现审计。
原文摘要 · Abstract (English)
Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,565 tweets annotated by three native-speaker annotators across eight batches and examine why inter-annotator agreement (IAA) declines over time. Despite an aggregate Randolph's free-marginal Kappa of $κ= 0.76$, "excellent," per-batch $κ$ falls by more than 32 points across the annotation task. Through six targeted analyses, we find that (i) label confusion concentrates on the negative/neutral boundary, (ii) two annotators show run-length drift consistent with autopilot labeling, and (iii) the dominant predictor of $κ$ is temporal simultaneity: tweets labeled within one minute achieve $κ= 0.98$, while those labeled more than a day apart reach only $κ= 0.65$. Annotation speed and tweet-level linguistic features show no meaningful association with $κ$. We benchmark three open multilingual encoders and proprietary models (GPT-5 and Gemini) on three-class sentiment classification; fine-tuning yields gains of 29 to 43 macro-F1 points over pretrained baselines, with GPT-5 few-shot leading overall (62.2 macro-F1). We release the dataset, per-annotation timestamps, and analysis code to support reproducible quality auditing for future African language NLP resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。