LLM生成的攻击叙事会加剧对心理疾病群体的污名化。
Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
- 构建网络框架分析语言模型生成攻击的传播路径
- 心理疾病相关实体在攻击网络中中心度极高(p值=4.06e-10)
- 适合关注AI伦理与社会影响的研究者阅读
大型语言模型(LLMs)已被证明对某些群体存在偏见。然而,针对易受伤害人群的无端攻击研究仍不充分。本文提出三项新贡献:(1)对心理疾病群体的攻击性文本进行显式评估;(2)提出基于网络的框架以研究偏见传播机制;(3)评估由此产生的污名化程度。通过对近期发布的大规模偏见审计数据集分析发现,心理疾病相关实体在攻击叙事网络中占据核心位置,其接近中心性显著更高(p值=4.06e-10),且聚类密度高(吉尼系数=0.7)。基于成熟污名化框架,生成链中针对心理健康障碍目标的标签成分显著增加。这些结果揭示了大模型在结构上倾向于强化有害言论,凸显了制定有效缓解策略的必要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been shown to demonstrate imbalanced biases against certain groups. However, the study of unprovoked targeted attacks by LLMs towards at-risk populations remains underexplored. Our paper presents three novel contributions: (1) the explicit evaluation of LLM-generated attacks on highly vulnerable mental health groups; (2) a network-based framework to study the propagation of relative biases; and (3) an assessment of the relative degree of stigmatization that emerges from these attacks. Our analysis of a recently released large-scale bias audit dataset reveals that mental health entities occupy central positions within attack narrative networks, as revealed by a significantly higher mean centrality of closeness (p-value = 4.06e-10) and dense clustering (Gini coefficient = 0.7). Drawing from an established stigmatization framework, our analysis indicates increased labeling components for mental health disorder-related targets relative to initial targets in generation chains. Taken together, these insights shed light on the structural predilections of large language models to heighten harmful discourse and highlight the need for suitable approaches for mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。