多智能体协作分析代码漏洞,精准识别并提升检测可靠性。
MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement
- 设计多智能体系统,通过交互式反馈实现跨过程代码理解。
- 在配对数据集上准确率比现有系统高62%以上,平均性能提升超600%。
- 适合安全研究者与开发团队用于高精度漏洞检测与评估。
开源软件的广泛使用带来了漏洞风险,现有漏洞检测方法受限于上下文理解不足、单轮交互和粗粒度评估,导致性能不佳与评估偏差。为此,我们提出MAVUL,一种结合上下文推理与交互优化的多智能体漏洞检测系统。其中,漏洞分析代理利用工具调用与上下文推理能力,实现跨过程代码理解,有效挖掘漏洞模式;通过跨角色智能体间的迭代反馈与决策优化,提升推理可靠性与预测准确性。此外,引入多维真实标签进行细粒度评估,显著提高评估精度与可信度。在配对漏洞数据集上的大量实验表明,MAVUL相比现有多智能体系统,配对准确率提升超过62%,相比单智能体系统平均性能提升超过600%。随着漏洞分析代理与安全架构代理间通信轮次增加,系统效果持续增强,凸显上下文推理在追踪漏洞传播中的重要性及反馈机制的关键作用。集成的评估代理作为无偏评判者,避免了误导性二元对比,更真实反映系统实际应用价值。
原文摘要 · Abstract (English)
The widespread adoption of open-source software (OSS) necessitates the mitigation of vulnerability risks. Most vulnerability detection (VD) methods are limited by inadequate contextual understanding, restrictive single-round interactions, and coarse-grained evaluations, resulting in undesired model performance and biased evaluation results. To address these challenges, we propose MAVUL, a novel multi-agent VD system that integrates contextual reasoning and interactive refinement. Specifically, a vulnerability analyst agent is designed to flexibly leverage tool-using capabilities and contextual reasoning to achieve cross-procedural code understanding and effectively mine vulnerability patterns. Through iterative feedback and refined decision-making within cross-role agent interactions, the system achieves reliable reasoning and vulnerability prediction. Furthermore, MAVUL introduces multi-dimensional ground truth information for fine-grained evaluation, thereby enhancing evaluation accuracy and reliability. Extensive experiments conducted on a pairwise vulnerability dataset demonstrate MAVUL's superior performance. Our findings indicate that MAVUL significantly outperforms existing multi-agent systems with over 62% higher pairwise accuracy and single-agent systems with over 600% higher average performance. The system's effectiveness is markedly improved with increased communication rounds between the vulnerability analyst agent and the security architect agent, underscoring the importance of contextual reasoning in tracing vulnerability flows and the crucial feedback role. Additionally, the integrated evaluation agent serves as a critical, unbiased judge, ensuring a more accurate and reliable estimation of the system's real-world applicability by preventing misleading binary comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。