通过分析标注员状态提升用户行为预测准确率
Improving User Behavior Prediction: Leveraging Annotator Metadata in Supervised Machine Learning Models
- 融合标注员疲劳、速度等元特征构建加权集成模型
- 在保留数据上性能提升14%,另一数据集上提升12%
- 高学历标注员更稳定高效,适合构建高质量标注体系
监督学习模型在从对话文本预测用户行为时表现不佳,主要受制于众包标签质量差和自然语言处理任务准确率低。本文提出元特征敏感加权编码集成模型(MSWEEM),整合标注员的疲劳、速度等元特征。实验表明,MSWEEM在保留数据上比标准集成模型性能提升14%,在另一数据集上提升12%。同时发现,引入标注员行为信号如速度与疲劳能显著提升模型表现。此外,具有硕士及以上资质的标注员标注更快速且一致性更高。随着标注质量不确定性加剧,理解标注员行为模式对提升用户行为预测准确性至关重要。
原文摘要 · Abstract (English)
Supervised machine-learning models often underperform in predicting user behaviors from conversational text, hindered by poor crowdsourced label quality and low NLP task accuracy. We introduce the Metadata-Sensitive Weighted-Encoding Ensemble Model (MSWEEM), which integrates annotator meta-features like fatigue and speeding. First, our results show MSWEEM outperforms standard ensembles by 14% on held-out data and 12% on an alternative dataset. Second, we find that incorporating signals of annotator behavior, such as speed and fatigue, significantly boosts model performance. Third, we find that annotators with higher qualifications, such as Master's, deliver more consistent and faster annotations. Given the increasing uncertainty over annotation quality, our experiments show that understanding annotator patterns is crucial for enhancing model accuracy in user behavior prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。