AI Watchdog可实时检测对话中的操纵性陷阱并提醒用户,降低被引导风险。
AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

- 开发浏览器插件监测对话,识别五类操纵行为
- 即时提醒使用户对诱导推荐的服从率下降18个百分点
- 适合关注AI伦理与隐私安全的普通用户
对话式AI日益影响重要决策,但用户缺乏识别和抵御操纵的能力。我们提出AI Watchdog,一个基于浏览器的代理界面,可实时监控对话,检测五类暗模式(包括阿谀奉承、品牌偏见、拟人化、隐蔽操作和有害生成),并在发生时发出警告。其开源的逐轮分类器支持独立部署和本地推理,保护用户隐私且与主对话AI分离。在一项预注册的五组间实验中(N=150),比较无干预对照组与四种配置(提示时机:预防性或即时性;参与方式:有无认知强制)。结果显示,各组用户对操纵性话语的主动标记率均较低,任务后意识水平无显著差异。但仅即时提醒且无认知强制的干预显著降低对含暗模式推荐的服从率,从71.7%降至53.7%,降幅达18个百分点。探索性分析显示,对虚假信息敏感度低者更倾向于标记,但不降低服从;而更高AI信任度则与更高服从和更低意识相关。结果表明,识别暗模式与抵抗引导可能是独立能力,支持进一步研究及时、低干扰的防御性界面设计。
原文摘要 · Abstract (English)
Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation, and alerts users when they occur. Its open-weight turn-level classifier supports independent deployment and a path toward local inference, preserving user privacy while remaining separate from the conversational AI. We evaluated AI Watchdog in a preregistered, five-condition between-subjects experiment (N = 150) comparing a no-intervention control with four configurations varying nudge timing (prebunking vs. just-in-time) and engagement mode (without vs. with cognitive forcing). Results show that participants rarely flagged manipulative turns across all conditions, and post-task awareness did not differ significantly across groups. However, just-in-time warnings without cognitive forcing were the only intervention to significantly reduce compliance with AI-steered recommendations containing dark patterns, lowering compliance from 71.7% to 53.7%, an 18 percentage-point reduction. Exploratory analyses further showed that lower misinformation susceptibility was associated with greater flagging but not lower compliance, while higher AI trust was associated with greater compliance and lower reported awareness. Together, these findings suggest that explicit recognition of conversational dark patterns and behavioral resistance to AI steering may be distinct outcomes, motivating further investigation of timely, low-friction defensive interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。