重新审视自然语言处理中的噪声,发现其可能蕴含重要社会信息。
Revisiting Noise in Natural Language Processing for Computational Social Science
- 通过多个案例研究分析文本噪声的来源与表现
- 指出部分噪声反映个体沟通风格与文化特征
- 适合关注社会科学研究中数据质量与语义理解的学者
计算社会科学(CSS)因海量人类生成内容而快速发展,但其研究面临独特挑战:理论主观性强,文本数据复杂且无结构。其中,噪声问题长期未受充分关注。本文通过一系列相互关联的案例研究,考察历史文献经光学字符识别产生的字符级错误、古语表达、主观任务标注不一致,以及大模型生成内容引入的噪声与偏见。研究挑战了噪声必然有害的传统观点,认为某些噪声可编码有意义信息,如个体交流风格或文化依赖性特征。文章强调应对噪声需具区分性策略,不同噪声类型应采用不同处理方式。
原文摘要 · Abstract (English)
Computational Social Science (CSS) is an emerging field driven by the unprecedented availability of human-generated content for researchers. This field, however, presents a unique set of challenges due to the nature of the theories and datasets it explores, including highly subjective tasks and complex, unstructured textual corpora. Among these challenges, one of the less well-studied topics is the pervasive presence of noise. This thesis aims to address this gap in the literature by presenting a series of interconnected case studies that examine different manifestations of noise in CSS. These include character-level errors following the OCR processing of historical records, archaic language, inconsistencies in annotations for subjective and ambiguous tasks, and even noise and biases introduced by large language models during content generation. This thesis challenges the conventional notion that noise in CSS is inherently harmful or useless. Rather, it argues that certain forms of noise can encode meaningful information that is invaluable for advancing CSS research, such as the unique communication styles of individuals or the culture-dependent nature of datasets and tasks. Further, this thesis highlights the importance of nuance in dealing with noise and the considerations CSS researchers must address when encountering it, demonstrating that different types of noise require distinct strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。