探究大模型在程序验证中的抽象推理失效机制
Understanding Formal Reasoning Failures in LLMs as Abstract Interpreters
- 设计新提示策略激发大模型的抽象解释能力
- 在22个SV-COMP程序上测试,发现推理错误模式
- 为大模型程序验证研究提供新方向
大型语言模型(LLMs)越来越多地用于程序验证,但对其在该过程中如何理解程序语义仍知之甚少。本文聚焦于基于抽象解释的不变量生成推理,提出两种新型提示策略,旨在引导大模型展现出此类推理能力。我们在多个最先进的大模型上,对来自广泛用于软件验证的SV-COMP基准套件的22个程序进行了评估。分析了生成不变量的正确性以及模型推理错误的关键主题模式。本工作旨在揭示大模型与程序验证交叉领域的全新研究机遇,推动大模型在验证任务中的应用及其推理能力的发展。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for program verification, and yet little is known about \emph{how} they reason about program semantics during this process. In this work, we focus on abstract interpretation based-reasoning for invariant generation and introduce two novel prompting strategies that aim to elicit such reasoning from LLMs. We evaluate these strategies across several state-of-the-art LLMs on 22 programs from the SV-COMP benchmark suite widely used in software verification. We analyze both the soundness of the generated invariants and the key thematic patterns in the models' reasoning errors. This work aims to highlight new research opportunities at the intersection of LLMs and program verification for applying LLMs to verification tasks and advancing their reasoning capabilities in this application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。