arXiv:2603.26898cs.CLcs.AI2026-03被引 1

LLM政治文本标注效果受多种因素影响,盲目遵循‘最佳实践’可能适得其反。

Magic Words or Methodical Work? Challenging Conventional Wisdom in LLM-Based Political Text Annotation

  • 控制变量对比六种开源模型在四类任务中的表现
  • 模型大小与性能无直接关系,跨族效率差异大
  • 常见提示工程技巧反而降低标注效果,需验证先行

政治科学家正快速采用大语言模型(LLMs)进行文本标注,但标注结果对实施方式的敏感性仍不清楚。多数评估仅测试单一模型或配置;模型选择、规模、学习方法与提示风格的交互作用,以及流行“最佳实践”是否经得起对照检验,尚不明确。本文在相同量化、硬件和提示模板条件下,对六种开源模型在四项政治科学标注任务中进行了受控评估。核心发现为方法论层面:交互效应主导主效应,看似合理的流程选择可能成为研究者自由度的来源。无单一模型、提示风格或学习方法始终最优,最佳模型随任务变化。两个推论随之而来:第一,模型规模不能可靠预测成本或性能——跨家族效率差异极大,某些大模型比小模型更省资源,同族中等规模模型常优于更大版本;第二,广泛推荐的提示工程技术对标注性能影响不一致,甚至产生负向效果。基于此基准结果,我们提出验证优先框架——包括流程决策的合理顺序、提示冻结建议、留出评估、报告标准及开源工具,帮助研究者透明地应对决策空间。

原文摘要 · Abstract (English)

Political scientists are rapidly adopting large language models (LLMs) for text annotation, yet the sensitivity of annotation results to implementation choices remains poorly understood. Most evaluations test a single model or configuration; how model choice, model size, learning approach, and prompt style interact, and whether popular "best practices" survive controlled comparison, are largely unexplored. We present a controlled evaluation of these pipeline choices, testing six open-weight models across four political science annotation tasks under identical quantisation, hardware, and prompt-template conditions. Our central finding is methodological: interaction effects dominate main effects, so seemingly reasonable pipeline choices can become consequential researcher degrees of freedom. No single model, prompt style, or learning approach is uniformly superior, and the best-performing model varies across tasks. Two corollaries follow. First, model size is an unreliable guide both to cost and to performance: cross-family efficiency differences are so large that some larger models are less resource-intensive than much smaller alternatives, while within model families mid-range variants often match or exceed larger counterparts. Second, widely recommended prompt engineering techniques yield inconsistent and sometimes negative effects on annotation performance. We use these benchmark results to develop a validation-first framework - with a principled ordering of pipeline decisions, guidance on prompt freezing and held-out evaluation, reporting standards, and open-source tools - to help researchers navigate this decision space transparently.

LLM标注政治文本方法论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。