用大模型模拟安全专家问卷响应,发现其结果有偏差,不宜直接替代真实专家。
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

- 对比多个大模型在不同提示下的生成答案,评估其可靠性。
- 模型回答一致性高但与真实专家差异明显,存在平均化和偏差问题。
- 适合用于预研和假设生成,不适用于正式专家调研结论。
专家问卷广泛用于安全研究中分析从业者工作流程与决策行为,但招募安全运营中心(SOC)专家因工作负荷重、易倦怠及保密限制,常面临样本量小的难题。大语言模型(LLMs)可大规模生成合成响应,是潜在替代方案,但缺乏可靠使用指导。本文提出一种方法论框架,评估LLM作为专家问卷替代或补充的适用性。基于实际SOC人员的回答,我们比较了多种模型在不同提示设置下的人格化与聚合型生成结果,衡量其稳定性、模型间一致性及与人类回答的对齐度。结果显示,尽管LLM生成内容内部一致,但系统性偏离真实专家,表现为方差降低、集中趋势偏差以及观点同质化。本研究为安全研究者提供了实证依据与实践建议:LLM适用于初步探索与假设生成,但不应取代真实专家访谈。
原文摘要 · Abstract (English)
Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。