揭露AI在漏洞评估中生成虚假信息的问题,提出可信验证新方法。
AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation
- 构建AI错误类型分类体系,定位推理机制缺陷
- 提出可量化的演绎覆盖度评分,揭示当前模型无法完全弥补推理差距
- 倡导主动神经符号验证,适合安全系统开发者与评估研究者
大型语言模型(LLMs)被广泛应用于网络安全漏洞评估,但其生成的“AI垃圾”——包括虚构漏洞、错误补丁及语义重构的漏洞报告——已引发信任危机。这些伪信息使人工排查负担剧增,如同拒绝服务攻击。本文通过系统文献综述,建立AI垃圾分类体系,揭示其根源:安全专家的因果演绎推理与当前LLM自回归概率生成之间的鸿沟。提出可测量的‘演绎覆盖度得分’作为代理指标,验证链式思考提示与工具调用代理虽能缩小差距却无法消除。分析现有缓解策略,指出被动检测与水印仅关注来源而非正确性,受熵增限制。主张采用主动神经符号验证,将每个组件映射至已有安全输入边界明确的系统。最后提出两个评估工具:CVE-Bench与Slop-Score,包含数据集构建、评分公式及防作弊设计。通过将评估标准从语言流畅性转向数学可验证性,为下一代可信AI漏洞评估系统提供路线图。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) into cybersecurity has transformed vulnerability assessment, but it has also produced a trustworthiness crisis driven by the unchecked proliferation of "AI slop." These artifacts, hallucinated vulnerabilities, plausible but incorrect patches, and semantically repackaged bug reports, impose a cognitive burden on human triage pipelines that mirrors a denial-of-service attack. This paper surveys the empirical evidence, identifies a unifying mechanism, and traces a path toward trustworthy triage. We formalize a taxonomy of AI slop grounded in a structured literature review and dissect its root cause: the gap between the causal deductive reasoning of security experts and the autoregressive probabilistic generation of current LLMs. We operationalize this gap through a measurable proxy, the Deductive Coverage Score, and show that chain-of-thought prompting and tool-using agents narrow but do not close it. We review mitigation strategies and argue that passive detection and watermarking target provenance rather than correctness, facing fundamental entropy constraints. We instead advocate for active neuro-symbolic verification, mapping each pipeline component to prior systems with documented limits on security inputs. Finally, we specify two evaluation instruments, CVE-Bench and Slop-Score, including dataset construction, metric formulas, and anti-gaming provisions. By shifting evaluation from linguistic fluency to mathematical verifiability, this survey provides a roadmap for securing emerging AI-driven triage systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。