测试大模型能否模拟用户安全隐私行为,发现效果有限但有改进空间。
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- 用30项真实实验构建基准,评估模型模拟人类安全态度与行为的匹配度。
- 所有模型平均得分50-64,新模型不一定更优,表现波动大。
- 设定理性权衡隐私成本与收益时,部分行为测试得分超95,适合安全风险预判。
越来越多研究假设大型语言模型(LLM)代理可反映人们对安全与隐私(S&P)威胁的态度和行为。若属实,这些模拟可为产品部署前的风险预测提供可扩展方案。本文通过SP-ABCBench——一个源自经验证的人类受试研究的30项测试基准——检验这一假设。该基准以0-100分衡量模拟结果与人类研究在态度、行为、一致性三个维度上的对齐程度,分数越高表示越接近。我们评估了12个LLM、4种角色构建策略及2种提示方法,发现仍有巨大提升空间:所有模型平均得分介于50至64之间。更新、更大或更智能的模型并未表现出可靠优势,有时甚至更差。然而,某些配置可实现高对齐性:例如,在要求模型采用有限理性并权衡隐私成本与感知收益时,部分行为测试得分超过95。我们已公开SP-ABCBench,以支持后续方法的可复现评估。
原文摘要 · Abstract (English)
A growing body of research assumes that large language model (LLM) agents can serve as proxies for how people form attitudes toward and behave in response to security and privacy (S&P) threats. If correct, these simulations could offer a scalable way to forecast S&P risks in products prior to deployment. We interrogate this assumption using SP-ABCBench, a new benchmark of 30 tests derived from validated S&P human-subject studies, which measures alignment between simulations and human-subjects studies on a 0-100 ascending scale, where higher scores indicate better alignment across three dimensions: Attitude, Behavior, and Coherence. Evaluating twelve LLMs, four persona construction strategies, and two prompting methods, we found that there remains substantial room for improvement: all models score between 50 and 64 on average. Newer, bigger, and smarter models do not reliably do better and sometimes do worse. Some simulation configurations, however, do yield high alignment: e.g., with scores above 95 for some behavior tests when agents are prompted to apply bounded rationality and weigh privacy costs against perceived benefits. We release SP-ABCBench to enable reproducible evaluation as methods improve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。