用理论驱动的合成数据,减少社会议题分类对标注数据的依赖。
From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
- 将社科测量工具转化为理论指导的合成文本数据
- 政治话题分类只需少量标注数据,性能下降不显著
- 有理论依据的合成数据比盲目生成更有效,适合社科研究者
计算文本分类在多维度社会构念识别中极具挑战性。近期研究指出,合成训练数据可通过呈现这些构念在文本中的表现形式来提升分类效果。本文系统评估了理论驱动的合成数据在提升社会构念测量中的潜力。具体而言,我们探索如何将社会科学中已有的测量工具(如问卷量表、标注手册)转化为理论指导下的合成数据生成方法。通过两项研究——性别歧视与政治话题测量——评估合成数据对微调文本分类模型的价值。尽管性别歧视研究结果有限,但发现合成数据在政治话题分类中极为有效:仅小幅性能损失下即可替代大量标注数据。此外,带有概念信息的合成数据显著优于无理论指导的数据生成方式。
原文摘要 · Abstract (English)
Computational text classification is a challenging task, especially for multi-dimensional social constructs. Recently, there has been increasing discussion that synthetic training data could enhance classification by offering examples of how these constructs are represented in texts. In this paper, we systematically examine the potential of theory-driven synthetic training data for improving the measurement of social constructs. In particular, we explore how researchers can transfer established knowledge from measurement instruments in the social sciences, such as survey scales or annotation codebooks, into theory-driven generation of synthetic data. Using two studies on measuring sexism and political topics, we assess the added value of synthetic training data for fine-tuning text classification models. Although the results of the sexism study were less promising, our findings demonstrate that synthetic data can be highly effective in reducing the need for labeled data in political topic classification. With only a minimal drop in performance, synthetic data allows for substituting large amounts of labeled data. Furthermore, theory-driven synthetic data performed markedly better than data generated without conceptual information in mind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。