通过社区数据发现大模型漏洞,揭示心理攻击比技术漏洞更有效。
PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models
- 从线上社区持续采集漏洞数据,多维度评分评估风险。
- 78%分类准确率,跨模型漏洞迁移能力有限。
- 适合关注LLM安全与社会工程威胁的研究者和开发者。
静态基准无法捕捉线上论坛中由社区实验引发的大语言模型新漏洞。我们提出PrompTrend系统,跨平台收集漏洞数据并采用多维评分进行评估,架构支持可扩展监控。对2025年1月至5月期间从在线社区收集的198个漏洞进行横断面分析,并在九个商用模型上测试,结果表明:部分架构中高级能力与漏洞增加相关;心理攻击显著优于技术利用;平台动态影响攻击效果,呈现可量化的模型特异性模式。PrompTrend漏洞评估框架达到78%分类准确率,但漏洞跨模型迁移能力有限,表明有效的大语言模型安全需超越传统周期性评估,实施全面的社会-技术监控。研究挑战了能力提升即更安全的假设,确立社区驱动的心理操纵为当前语言模型的主要威胁向量。
原文摘要 · Abstract (English)
Static benchmarks fail to capture LLM vulnerabilities emerging through community experimentation in online forums. We present PrompTrend, a system that collects vulnerability data across platforms and evaluates them using multidimensional scoring, with an architecture designed for scalable monitoring. Cross-sectional analysis of 198 vulnerabilities collected from online communities over a five-month period (January-May 2025) and tested on nine commercial models reveals that advanced capabilities correlate with increased vulnerability in some architectures, psychological attacks significantly outperform technical exploits, and platform dynamics shape attack effectiveness with measurable model-specific patterns. The PrompTrend Vulnerability Assessment Framework achieves 78% classification accuracy while revealing limited cross-model transferability, demonstrating that effective LLM security requires comprehensive socio-technical monitoring beyond traditional periodic assessment. Our findings challenge the assumption that capability advancement improves security and establish community-driven psychological manipulation as the dominant threat vector for current language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。