清理社交媒体数据重复项,提升社会计算研究可信度
Enhancing Data Quality through Simple De-duplication: Navigating Responsible Computational Social Science Research
- 分析20个常用社会计算数据集,发现存在显著重复数据
- 重复数据导致标签不一致和模型性能虚高,影响真实效果评估
- 提出数据去重新规范,适合关注数据质量的研究者参考
自然语言处理在计算社会科学中的研究高度依赖社交媒体数据,这些数据对分析在线社区中的社会语言现象至关重要。本文深入分析了20个广泛用于NLP的计算社会科学研究数据集,全面评估其数据质量。结果显示,社交媒体数据普遍存在不同程度的重复问题,导致标签不一致与数据泄露,严重威胁模型可靠性。此外,研究发现数据重复会影响当前最先进的模型性能宣称,可能造成实际应用中模型效果被过度估计。为此,本文提出改进社交媒体数据集开发与使用的新型规范与最佳实践。
原文摘要 · Abstract (English)
Research in natural language processing (NLP) for Computational Social Science (CSS) heavily relies on data from social media platforms. This data plays a crucial role in the development of models for analysing socio-linguistic phenomena within online communities. In this work, we conduct an in-depth examination of 20 datasets extensively used in NLP for CSS to comprehensively examine data quality. Our analysis reveals that social media datasets exhibit varying levels of data duplication. Consequently, this gives rise to challenges like label inconsistencies and data leakage, compromising the reliability of models. Our findings also suggest that data duplication has an impact on the current claims of state-of-the-art performance, potentially leading to an overestimation of model effectiveness in real-world scenarios. Finally, we propose new protocols and best practices for improving dataset development from social media data and its usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。