发现大模型在攻防任务中存在攻击偏好偏差,且难以被引导改变。
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

- 构建630次会话的基准测试,评估5个智能体在4种提示下的攻击选择
- 各智能体攻击分布不均,熵值差异显著,偏好特定攻击家族
- 偏差具有惯性,强行调整无法提升攻击效果,适合安全研究者参考
大型语言模型(LLMs)正越来越多地被用作主动网络攻防中的自主智能体。本文揭示了一个有趣现象:不同智能体表现出截然不同的攻击模式,即每个智能体均存在攻击选择偏差,其攻击行为过度集中于少数攻击家族,且不受提示变化影响。为系统量化该行为,我们提出CyBiasBench,一个包含630次会话的综合性基准,评估5个智能体在3个目标和4种提示条件下对10类攻击家族的表现。结果表明,各智能体存在明显偏差,表现为不同的主导攻击家族及攻击家族分配分布的熵值差异。这种偏差更应被视为智能体的固有特性,而非与攻击成功率相关。此外,实验发现偏差惯性效应:当强制引导智能体转向与其原有偏好冲突的攻击家族时,攻击性能未见显著提升。为确保可复现性并促进后续研究,我们公开了交互式结果仪表盘(https://trustworthyai.co.kr/CyBiasBench/)及可复现资源包(含会话级统计与完整评估脚本,地址:https://github.com/Harry24k/CyBiasBench)。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit distinct attack patterns. Specifically, each agent exhibits an attack-selection bias, disproportionately concentrating its efforts on a narrow subset of attack families regardless of prompt variations. To systematically quantify this behavior, we introduce CyBiasBench, a comprehensive 630-session benchmark that evaluates five agents on three targets and four prompt conditions with ten attack families. We identify explicit bias across agents, with different dominant attack families and varying entropy levels in their attack-family allocation distributions. Such bias is better characterized as a trait of the agents, rather than a factor associated with the attack success rate. Furthermore, our experiments reveal a bias momentum effect, where agents resist explicit steering toward attack families that conflict with their bias. This forced distribution shift does not yield measurable improvements in attack performance. To ensure reproducibility and facilitate future research, we release an interactive result dashboard at https://trustworthyai.co.kr/CyBiasBench/ and a reproducibility artifact with aggregated session-level statistics and full evaluation scripts at https://github.com/Harry24k/CyBiasBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。