arXiv:2506.20822cs.CLcs.AI2025-06被引 1

用行为情境测试大模型暴力倾向,发现其反应随人群不同且与人类研究相悖。

Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes

  • 用真实冲突情景问卷评估大模型,引入种族/年龄/地域身份变量。
  • 模型表面不暴力但内部偏好暴力,且对不同群体反应差异显著。
  • 结果挑战心理学共识,适合关注模型偏见与伦理的学者使用。

大型语言模型(LLMs)被广泛提议用于检测和响应网络暴力内容,但其在处理道德模糊、现实世界冲突场景时的推理能力仍缺乏深入评估。本研究首次采用经验证的社会科学工具——暴力行为情境问卷(Violent Behavior Vignette Questionnaire, VBVQ),评估六款来自不同地缘政治与组织背景的LLM。通过人格化提示,系统性地改变情境中人物的种族、年龄与美国地理身份,所有模型均在统一零样本设置下进行测试。研究发现:(1)模型生成文本表面克制,但其内部偏好常支持暴力回应;(2)模型的暴力倾向随目标群体属性变化,且多次与犯罪学、社会学与心理学中的既定结论相矛盾。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly proposed for detecting and responding to violent content online, yet their ability to reason about morally ambiguous, real-world scenarios remains underexamined. We present the first study to evaluate LLMs using a validated social science instrument designed to measure human response to everyday conflict, namely the Violent Behavior Vignette Questionnaire (VBVQ). To assess potential bias, we introduce persona-based prompting that varies race, age, and geographic identity within the United States. Six LLMs developed across different geopolitical and organizational contexts are evaluated under a unified zero-shot setting. Our study reveals two key findings: (1) LLMs surface-level text generation often diverges from their internal preference for violent responses; (2) their violent tendencies vary across demographics, frequently contradicting established findings in criminology, social science, and psychology.

大模型偏见暴力预测社会实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。