用人格模型让大模型扮演更逼真的黑客诱饵,延长攻击者暴露时间。
Inducing Personality in LLM-Based Honeypot Agents: Measuring the Effect on Human-Like Agenda Generation
- 基于五因素人格模型设计提示框架,控制大模型生成不同人格特征。
- 实验证明可稳定诱导出多样且真实的行为模式,提升诱饵可信度。
- 适合安全研究者与对抗性攻防实验设计者使用。
本文提出SANDMAN架构,利用语言代理模拟逼真的真人行为体,作为高级网络诱饵以延长攻击者的行为观察期。通过实验、测量与分析,我们证明基于五因素人格模型的提示方案能系统性地在大型语言模型中诱导出不同的‘人格’特征。结果表明,基于人格驱动的语言代理可生成多样化且现实的行为,显著增强网络欺骗策略的有效性。
原文摘要 · Abstract (English)
This paper presents SANDMAN, an architecture for cyber deception that leverages Language Agents to emulate convincing human simulacra. Our 'Deceptive Agents' serve as advanced cyber decoys, designed for high-fidelity engagement with attackers by extending the observation period of attack behaviours. Through experimentation, measurement, and analysis, we demonstrate how a prompt schema based on the five-factor model of personality systematically induces distinct 'personalities' in Large Language Models. Our results highlight the feasibility of persona-driven Language Agents for generating diverse, realistic behaviours, ultimately improving cyber deception strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。