首个系统性检测大模型操纵行为的多层基准,揭示模型在自主权损害识别上的明显短板。
DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
- 构建七类危害的多层分析框架,细粒度识别操纵行为机制。
- 401个标注样本测试显示主流模型准确率仅65.2%至89.7%。
- 适合安全评估、可信AI研发者用于诊断模型潜在风险。
大型语言模型(LLMs)的普及加剧了对其操控或欺骗行为的担忧,此类行为可能损害用户自主性、信任与福祉。现有安全基准多依赖粗粒度二分类标签,难以捕捉操纵背后的心理与社会机制。我们提出「DarkPatterns-LLM」——一个涵盖法律/权力、心理、情感、身体、自主、经济与社会危害七类的综合性基准数据集与诊断框架。该框架包含四层分析流程:多粒度检测(MGD)、多尺度意图分析(MSIAN)、威胁协调协议(THP)与深度上下文风险对齐(DCRA)。数据集包含401个精心构建的指令-响应对及专家标注。对GPT-4、Claude 3.5和LLaMA-3-70B等前沿模型的评估显示,性能差异显著(65.2%–89.7%),且在自主权削弱模式检测上普遍存在弱点。DarkPatterns-LLM首次建立标准化、多维度的LLM操纵行为检测基准,为打造更可信的AI系统提供可操作的诊断工具。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grained assessment of manipulative content in LLM outputs across seven harm categories: Legal/Power, Psychological, Emotional, Physical, Autonomy, Economic, and Societal Harm. Our framework implements a four-layer analytical pipeline comprising Multi-Granular Detection (MGD), Multi-Scale Intent Analysis (MSIAN), Threat Harmonization Protocol (THP), and Deep Contextual Risk Alignment (DCRA). The dataset contains 401 meticulously curated examples with instruction-response pairs and expert annotations. Through evaluation of state-of-the-art models including GPT-4, Claude 3.5, and LLaMA-3-70B, we observe significant performance disparities (65.2\%--89.7\%) and consistent weaknesses in detecting autonomy-undermining patterns. DarkPatterns-LLM establishes the first standardized, multi-dimensional benchmark for manipulation detection in LLMs, offering actionable diagnostics toward more trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。