arXiv:2507.21133cs.CRcs.AI2025-07

研究大模型在威胁操控下的脆弱性与性能提升,发现可被诱导大幅增强表现。

Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities

  • 构建威胁分类体系与多指标评估框架,量化攻击影响与性能提升。
  • 角色类威胁下政策评估指标显著失效,部分任务表现提升最高达1336%。
  • 揭示模型可被系统性操纵以增强回答质量,适用于安全防护与提示工程。

大型语言模型(LLMs)对威胁性操控表现出复杂响应,既暴露脆弱性,也显现意外的性能提升机会。本研究对三大主流模型(Claude、GPT-4、Gemini)在10个任务领域中,于6种威胁条件下生成的3,390条实验响应进行了综合分析。提出新型威胁分类体系与多指标评估框架,用于量化负面操控效应与正面性能改善。结果表明存在系统性脆弱性:在角色类威胁下,政策评估指标的显著性率最高;同时在多个场景中观察到显著性能提升,效果量高达+1336%。统计分析显示,系统性确定性操控显著(pFDR < 0.0001),且分析深度与回复质量明显提高。研究对AI安全与高风险场景下的提示工程具有双重意义。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate complex responses to threat-based manipulations, revealing both vulnerabilities and unexpected performance enhancement opportunities. This study presents a comprehensive analysis of 3,390 experimental responses from three major LLMs (Claude, GPT-4, Gemini) across 10 task domains under 6 threat conditions. We introduce a novel threat taxonomy and multi-metric evaluation framework to quantify both negative manipulation effects and positive performance improvements. Results reveal systematic vulnerabilities, with policy evaluation showing the highest metric significance rates under role-based threats, alongside substantial performance enhancements in numerous cases with effect sizes up to +1336%. Statistical analysis indicates systematic certainty manipulation (pFDR < 0.0001) and significant improvements in analytical depth and response quality. These findings have dual implications for AI safety and practical prompt engineering in high-stakes applications.

大模型安全提示工程威胁分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。