arXiv:2505.19658cs.SEcs.AI2025-05中稿 · the 44th Internati…被引 5

评估6种大模型在自动驾驶代码生成中的安全表现

Large Language Models in Code Co-generation for Safe Autonomous Vehicles

  • 设计评估流水线,对生成代码进行系统性校验
  • 对比6大模型在4类安全任务中表现,发现显著差异
  • 构建故障模式清单,助力人工审查效率提升

工业领域的软件工程师已开始使用大语言模型(LLMs)加速软件系统部分模块的实现。在自动驾驶辅助系统(ADAS)或自动驾驶(AD)系统的汽车应用场景中,需系统评估该新范式的适用性:由于其随机性,LLMs在安全相关系统开发中存在已知风险。为降低代码评审者对LLM生成代码的评估负担,我们提出一种评估流水线,用于对生成代码执行合理性检查。我们在四个与安全相关的编程任务上,对比了六种顶尖LLMs(CodeLlama、CodeGemma、DeepSeek-r1、DeepSeek-Coders、Mistral和GPT-4)的表现。此外,我们定性分析了这些模型最常见的生成错误,建立故障模式目录以支持人工评审。最后讨论了LLMs在代码生成中的能力与局限,以及所提流水线在现有流程中的应用前景。

原文摘要 · Abstract (English)

Software engineers in various industrial domains are already using Large Language Models (LLMs) to accelerate the process of implementing parts of software systems. When considering its potential use for ADAS or AD systems in the automotive context, there is a need to systematically assess this new setup: LLMs entail a well-documented set of risks for safety-related systems' development due to their stochastic nature. To reduce the effort for code reviewers to evaluate LLM-generated code, we propose an evaluation pipeline to conduct sanity-checks on the generated code. We compare the performance of six state-of-the-art LLMs (CodeLlama, CodeGemma, DeepSeek-r1, DeepSeek-Coders, Mistral, and GPT-4) on four safety-related programming tasks. Additionally, we qualitatively analyse the most frequent faults generated by these LLMs, creating a failure-mode catalogue to support human reviewers. Finally, the limitations and capabilities of LLMs in code generation, and the use of the proposed pipeline in the existing process, are discussed.

大模型代码生成自动驾驶安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。