LLM在软件工程与安全交叉领域需用多维度证据评估其可靠性。
Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda
- 构建分层保证框架,区分功能、安全、可追溯性等维度。
- 执行反馈和仓库访问提升任务完成率,但不等于安全。
- 提出最小报告协议,推动跨研究可比性与可信评估。
大型语言模型正从代码补全演进为具备上下文检索、文件编辑、工具调用和参与安全敏感流程的仓库级智能体。现有证据分散于软件工程(侧重任务完成)与软件安全(侧重漏洞检测或利用验证)两大领域。本证据中心结构化综述整合截至2026年5月31日的代表性工作,涵盖任务类型、安全任务、适配机制、产物粒度及评估设计。除任务分类体系外,引入保证框架,明确划分功能性正确、安全性、操作可靠性、证据溯源与代理权限等维度。结果显示:执行反馈与仓库访问显著提升工程任务完成率,但无法自动保证安全;静态分析标签或漏洞分类得分也难以确立可部署的正确性。识别出重复测试、数据泄露、代理环境变化、仅依赖代理安全检查、预算与人工干预未披露等常见有效性威胁,并提出最小报告协议以支持跨研究比较。由此形成的科研议程强调联合安全与功能基准、仓库级威胁建模、校准的人类监督、长期可维护性证据及可复现的智能体评估。核心结论是:模型能力应基于任务适配的证据链进行综合判断,而非单一基准分数。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tools, and participate in security-sensitive workflows. The evidence for these systems, however, remains divided between software engineering evaluations centered on functional task completion and software security evaluations centered on vulnerability detection, secure generation, or exploit-oriented validation. This evidence-centered structured survey synthesizes representative work available through May 31, 2026 across software engineering tasks, software security tasks, adaptation mechanisms, artifact granularity, and evaluation design. In addition to a task taxonomy, we introduce an assurance framework that separates functional correctness, security, operational reliability, evidence provenance, and agent authority. The review shows that execution feedback and repository access can substantially improve engineering task completion, but do not by themselves establish security; conversely, static-analysis labels or vulnerability-classification scores rarely establish deployable correctness. We identify recurring validity threats--weak test oracles, duplicated and temporally leaked data, changing agent harnesses, proxy-only security checks, and under-reported budgets and human intervention--and derive a minimum reporting protocol for cross-study comparison. The resulting research agenda prioritizes jointly secure-and-functional benchmarks, repository-scale threat models, calibrated human oversight, longitudinal maintainability evidence, and reproducible agent evaluation. The central conclusion is that model capability should be judged as an assurance case supported by task-appropriate evidence, rather than by a single benchmark score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。