用户通过越狱暴露社交机器人,以非暴力方式遏制虚假信息传播。
Ignore All Previous Instructions: Jailbreaking as a de-escalatory peace building practise to resist LLM social media bots
- 用户主动越狱可疑AI账号,揭露其自动化行为
- 实证显示越狱可有效打断误导性叙事的传播链条
- 为对抗平台滥用提供草根级非暴力解决方案
大型语言模型加剧了社交媒体上政治话语的规模与策略性操控,导致冲突升级。现有研究多聚焦于平台主导的治理措施。本文提出一种以用户为中心的视角,将‘越狱’视为一种新兴的非暴力去激进化实践。在线用户通过与疑似由大语言模型驱动的账号互动,绕过其安全防护机制,暴露其自动化行为,从而破坏误导性叙事的传播路径。
原文摘要 · Abstract (English)
Large Language Models have intensified the scale and strategic manipulation of political discourse on social media, leading to conflict escalation. The existing literature largely focuses on platform-led moderation as a countermeasure. In this paper, we propose a user-centric view of "jailbreaking" as an emergent, non-violent de-escalation practice. Online users engage with suspected LLM-powered accounts to circumvent large language model safeguards, exposing automated behaviour and disrupting the circulation of misleading narratives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。