arXiv:2507.21831cs.CLcs.AI2025-07被引 1

提出HALC框架,系统化找到大模型在社科编码中的最优提示策略。

Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences

  • 构建通用管道HALC,整合多种提示策略自动优化编码提示。
  • 在Mistral NeMo上实现单变量编码信度α=0.76~0.78,双变量α=0.71~0.74。
  • 适合需高可靠自动化编码的社科研究者,尤其关注提示工程优化。

大语言模型在任务自动化中广泛应用,包括社会科学中的自动化编码。然而,尽管已有多种提示策略,其效果在不同模型和任务间差异显著,仍普遍依赖试错。本文提出HALC——一个通用管道,可系统可靠地为任意编码任务和模型构建最优提示,支持集成任何相关提示策略。为验证该方法,我们向本地LLM发送了超过两百万次请求,共1,512个独立提示。基于少量专家编码(真实标签)评估性能,发现使用Mistral NeMo时,所提提示策略对单变量编码的信度达α=0.76(气候)、α=0.78(运动),双变量编码分别为α=0.71和α=0.74。提示设计旨在对齐代码本,而非优化代码本以适配模型。论文揭示了不同提示策略的有效性、关键影响因素,并识别出每项任务与模型组合下的可靠提示。

原文摘要 · Abstract (English)

LLMs are seeing widespread use for task automation, including automated coding in the social sciences. However, even though researchers have proposed different prompting strategies, their effectiveness varies across LLMs and tasks. Often trial and error practices are still widespread. We propose HALC$-$a general pipeline that allows for the systematic and reliable construction of optimal prompts for any given coding task and model, permitting the integration of any prompting strategy deemed relevant. To investigate LLM coding and validate our pipeline, we sent a total of 1,512 individual prompts to our local LLMs in over two million requests. We test prompting strategies and LLM task performance based on few expert codings (ground truth). When compared to these expert codings, we find prompts that code reliably for single variables ($α$climate = .76; $α$movement = .78) and across two variables ($α$climate = .71; $α$movement = .74) using the LLM Mistral NeMo. Our prompting strategies are set up in a way that aligns the LLM to our codebook$-$we are not optimizing our codebook for LLM friendliness. Our paper provides insights into the effectiveness of different prompting strategies, crucial influencing factors, and the identification of reliable prompts for each coding task and model.

提示工程自动化编码大模型应用社会科学研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。