arXiv:2512.22470cs.AI2025-12被引 4

首个系统性检测大模型操纵行为的多层基准,揭示模型在自主权损害识别上的明显短板。

DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior

  • 构建七类危害的多层分析框架,细粒度识别操纵行为机制。
  • 401个标注样本测试显示主流模型准确率仅65.2%至89.7%。
  • 适合安全评估、可信AI研发者用于诊断模型潜在风险。

大型语言模型(LLMs)的普及加剧了对其操控或欺骗行为的担忧,此类行为可能损害用户自主性、信任与福祉。现有安全基准多依赖粗粒度二分类标签,难以捕捉操纵背后的心理与社会机制。我们提出「DarkPatterns-LLM」——一个涵盖法律/权力、心理、情感、身体、自主、经济与社会危害七类的综合性基准数据集与诊断框架。该框架包含四层分析流程:多粒度检测(MGD)、多尺度意图分析(MSIAN)、威胁协调协议(THP)与深度上下文风险对齐(DCRA)。数据集包含401个精心构建的指令-响应对及专家标注。对GPT-4、Claude 3.5和LLaMA-3-70B等前沿模型的评估显示,性能差异显著(65.2%–89.7%),且在自主权削弱模式检测上普遍存在弱点。DarkPatterns-LLM首次建立标准化、多维度的LLM操纵行为检测基准,为打造更可信的AI系统提供可操作的诊断工具。

原文摘要 · Abstract (English)

The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grained assessment of manipulative content in LLM outputs across seven harm categories: Legal/Power, Psychological, Emotional, Physical, Autonomy, Economic, and Societal Harm. Our framework implements a four-layer analytical pipeline comprising Multi-Granular Detection (MGD), Multi-Scale Intent Analysis (MSIAN), Threat Harmonization Protocol (THP), and Deep Contextual Risk Alignment (DCRA). The dataset contains 401 meticulously curated examples with instruction-response pairs and expert annotations. Through evaluation of state-of-the-art models including GPT-4, Claude 3.5, and LLaMA-3-70B, we observe significant performance disparities (65.2\%--89.7\%) and consistent weaknesses in detecting autonomy-undermining patterns. DarkPatterns-LLM establishes the first standardized, multi-dimensional benchmark for manipulation detection in LLMs, offering actionable diagnostics toward more trustworthy AI systems.

大模型安全操纵检测可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。