arXiv:2605.07422cs.SEcs.AI2026-05中稿 · the 1st Internatio…被引 2

测试三种大模型在心理安全编码中的提示工程效果,发现多步提示提升稳定性。

Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study

论文配图:Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
图 1 · 摘自论文原文
  • 对比零样本与多步提示策略,评估大模型编码一致性。
  • 多步提示使Claude Haiku的评分一致率提升0.034(显著),其他模型无提升。
  • 所有模型高估负面反馈,低估担忧表达,适合研究者参考提示设计。

定性分析在理解软件工程中的人文与社会因素方面至关重要,但其过程依赖研究人员的主观判断,且对方法选择(如提示设计)敏感。近期大语言模型(LLMs)为支持此类分析提供了可能,但其在不同提示条件下复现人类定性推理的可靠性仍缺乏实证检验。本研究对Claude Haiku、DeepSeek-Chat和Gemini 2.5 Flash三款模型,采用零样本与多步封闭编码两种提示策略,在十次独立运行中以Cohen's kappa为关键一致性指标进行控制实验。结果表明,多步提示显著提升了Claude Haiku的一致性(Delta kappa = +0.034,Wilcoxon p = 0.004),但对DeepSeek-Chat和Gemini 2.5 Flash无效。模型内部稳定性差异明显:DeepSeek-Chat与Claude Haiku方差最低(标准差约0.017),而Gemini 2.5 Flash最不稳定(标准差=0.038)。所有模型均系统性高估“分享负面反馈”(偏差比高达5.25倍),同时持续低估“表达担忧”。研究为软件工程领域大模型辅助定性编码的提示工程提供了实证指导。

原文摘要 · Abstract (English)

Qualitative analysis plays a pivotal role in understanding the human and social aspects of software engineering. However, it remains a demanding process shaped by the subjective interpretation of individual researchers and sensitive to methodological choices such as prompt design. Recent advancements in Large Language Models (LLMs) offer promising opportunities to support this type of analysis, although their reliability in reproducing human qualitative reasoning under varying prompting conditions remains largely untested. This study presents a controlled empirical evaluation of three LLMs -- Claude Haiku, DeepSeek-Chat, and Gemini 2.5 Flash -- across two prompt engineering strategies (zero-shot and multi-shot closed coding), using Cohen's kappa as the primary agreement metric over ten independent runs per configuration. Results suggest that multi-shot prompting significantly improves agreement for Claude Haiku (Delta kappa = +0.034, Wilcoxon p = 0.004) but not for DeepSeek-Chat or Gemini 2.5 Flash. Intra-model stability varies substantially -- DeepSeek-Chat and Claude Haiku exhibit the lowest variance (SD approx. 0.017), while Gemini 2.5 Flash is the least stable (SD = 0.038). A systematic over-prediction of "Sharing Negative Feedback" is identified across all models (bias ratios up to 5.25x), alongside consistent under-prediction of "Expressing Concerns." Collectively, these findings provide empirical evidence for prompt engineering guidelines in LLM-assisted qualitative coding for software engineering research.

大模型提示工程定性分析软件工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。