arXiv:2608.15092cs.CRcs.AI2026-08

构建基准测试工具,量化大模型代码编辑中的安全风险漂移

WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing

论文配图:WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing
图 1 · 摘自论文原文
  • 设计连续风险表示法,整合多种漏洞信号
  • 在400个真实代码任务中发现显著安全风险变化
  • 适合关注AI代码生成安全性的研究者和开发者

本文提出WeSCE,一个用于量化弱安全约束下大模型代码编辑中安全漂移的基准。该基准包含400个可执行的真实世界代码程序,涵盖功能新增、删除、缺陷修复和重构任务。为衡量安全漂移,我们提出一种连续风险表示方法,通过统一框架聚合异构漏洞信号,并定义了反映整体风险、最坏情况严重性及漏洞分布变化的漂移度量,提供从平均行为到极端场景的多尺度安全视图。

原文摘要 · Abstract (English)

In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks specify only functional objectives without explicit security requirements. WeSCE consists of 400 executable programs derived from real-world code, covering feature addition, feature removal, bug fixing, and refactoring. To quantify security drift, we propose a continuous risk representation that aggregates heterogeneous vulnerability signals through a unified formulation, and define drift measures capturing changes in overall risk, worst-case severity, and vulnerability distribution under code transformations, providing a multi-scale view of security spanning average-case behavior to worst-case emphasis.

代码安全大模型评估安全漂移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。