LLM分类别忽视概念界定,会导致分析偏差,难以通过提升模型精度纠正。
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
- 强调在使用LLM前必须明确待分类概念的定义
- 模拟显示概念化错误会系统性扭曲结果,无法仅靠提升准确率修复
- 适合关注社会科学研究中模型应用可信度的学者
生成式大语言模型(LLMs)现广泛用于计算社会科学中的文本分类。本文关注提示前后的关键环节——概念界定与下游统计推断,指出这些步骤在多数LLM时代的社会科学研究中被忽视。我们主张,LLM可能诱使研究者跳过概念界定,从而引发概念化误差,导致下游估计偏差。通过模拟实验表明,此类偏差无法仅通过提高LLM准确率或事后校正方法消除。最后提醒研究者:在大模型时代,概念界定仍是首要关切,并提供了低成本、低偏差、低方差下游估计的具体建议。
原文摘要 · Abstract (English)
Generative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conceptualization of concepts to be classified and using LLM predictions in downstream statistical inference -- which we argue have been overlooked in much of LLM-era CSS. We claim LLMs can tempt analysts to skip the conceptualization step, creating conceptualization errors that bias downstream estimates. Using simulations, we show that this conceptualization-induced bias cannot be corrected for solely by increasing LLM accuracy or post-hoc bias correction methods. We conclude by reminding CSS analysts that conceptualization is still a first-order concern in the LLM-era and provide concrete advice on how to pursue low-cost, unbiased, low-variance downstream estimates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。