arXiv:2601.02941cs.CRcs.AI2026-01被引 4

构建真实SAST误报测试基准,助力自动化漏洞筛查

SastBench: A Benchmark for Testing Agentic SAST Triage

  • 用真实漏洞+过滤后误报构造测试集
  • 揭示现有智能代理在真实场景下的性能差距
  • 适合安全工具开发者与自动化研究者

SAST(静态应用安全测试)是防御性网络安全中最广泛使用的技术之一,被商业和非商业组织用于识别软件中的潜在漏洞。尽管其用途广泛,但会产生大量误报,需耗费人力进行手动筛选(即归类)。虽然大模型驱动的智能体在自动化安全任务中展现出潜力,但现有基准无法模拟真实世界中SAST发现结果的分布。本文提出SastBench,一个用于评估SAST归类智能体的基准,结合真实CVE作为真阳性,以及经过滤后的SAST工具输出作为近似假阳性。该基准设计具有智能体无关性。我们在该基准上评估多种智能体,展示性能对比分析,提供数据集详细剖析,并讨论对未来发展的启示。

原文摘要 · Abstract (English)

SAST (Static Application Security Testing) tools are among the most widely used techniques in defensive cybersecurity, employed by commercial and non-commercial organizations to identify potential vulnerabilities in software. Despite their great utility, they generate numerous false positives, requiring costly manual filtering (aka triage). While LLM-powered agents show promise for automating cybersecurity tasks, existing benchmarks fail to emulate real-world SAST finding distributions. We introduce SastBench, a benchmark for evaluating SAST triage agents that combines real CVEs as true positives with filtered SAST tool findings as approximate false positives. SastBench features an agent-agnostic design. We evaluate different agents on the benchmark and present a comparative analysis of their performance, provide a detailed analysis of the dataset, and discuss the implications for future development.

安全测试智能体漏洞检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。