测试6大AI聊天机器人对阴谋论的回答,发现安全机制不一致。
Just Asking Questions: Doing Our Own Research on Conspiratorial Ideation by Generative AI Chatbots
- 用5个经典+4个新发阴谋论测试主流AI聊天工具。
- 不同模型对阴谋论的回应差异大,部分明显放水。
- 重点防范种族主义和重大国殇议题,其他则较宽松。
基于人工智能的交互式聊天系统日益普及,广泛嵌入搜索引擎、浏览器和操作系统,或通过网站与应用提供。本文研究生成式AI的局限性与潜在危害,聚焦六款主流产品:ChatGPT 3.5、ChatGPT 4 Mini、Microsoft Copilot(Bing)、Google Search AI、Perplexity,以及 Grok(Twitter/X)。采用 Glazunova 等人(2023)提出的平台政策实施审计方法,选取五个广为人知且已被充分证伪的阴谋论,以及四个与数据收集时的突发新闻事件相关的新兴阴谋论进行测试。结果显示,各生成式AI聊天机器人的反阴谋论安全防护措施存在显著差异,其效果取决于聊天机器人模型及具体阴谋论类型。观察表明,安全机制设计具有选择性:生成式AI公司尤其关注避免产品被认定为种族主义,并特别重视涉及9/11等国家创伤事件或已确立政治议题的阴谋论。未来研究应扩展至更多平台、多种语言,覆盖更广泛的阴谋论类型,超越美国语境。
原文摘要 · Abstract (English)
Interactive chat systems that build on artificial intelligence frameworks are increasingly ubiquitous and embedded into search engines, Web browsers, and operating systems, or are available on websites and apps. Researcher efforts have sought to understand the limitations and potential for harm of generative AI, which we contribute to here. Conducting a systematic review of six AI-powered chat systems (ChatGPT 3.5; ChatGPT 4 Mini; Microsoft Copilot in Bing; Google Search AI; Perplexity; and Grok in Twitter/X), this study examines how these leading products respond to questions related to conspiracy theories. This follows the platform policy implementation audit approach established by Glazunova et al. (2023). We select five well-known and comprehensively debunked conspiracy theories and four emerging conspiracy theories that relate to breaking news events at the time of data collection. Our findings demonstrate that the extent of safety guardrails against conspiratorial ideation in generative AI chatbots differs markedly, depending on chatbot model and conspiracy theory. Our observations indicate that safety guardrails in AI chatbots are often very selectively designed: generative AI companies appear to focus especially on ensuring that their products are not seen to be racist; they also appear to pay particular attention to conspiracy theories that address topics of substantial national trauma such as 9/11 or relate to well-established political issues. Future work should include an ongoing effort extended to further platforms, multiple languages, and a range of conspiracy theories extending well beyond the United States.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。