首个评估大模型生成ASP代码能力的基准测试框架
BLAST: Benchmarking LLMs with ASP-based Structured Testing

- 基于答案集编程设计结构化测试方法
- 在8个主流大模型上验证10个图论问题的生成准确率
- 适合关注逻辑推理与程序生成的研究者
大型语言模型(LLMs)在自然语言理解、对话系统和代码生成等多个任务中展现出卓越性能。然而,目前对它们在命题式编程范式(如答案集编程,ASP)中的表现关注仍不足。本文提出BLAST:首个专用于评估大模型生成ASP代码准确性的基准测试方法及配套数据集。BLAST提供结构化评估框架,包含两项针对ASP代码生成的新语义度量指标。论文通过实证评估,考察了来自ASP文献的10个典型图相关问题,以及8个前沿大模型的表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable performance across a broad spectrum of tasks, including natural language understanding, dialogue systems, and code generation. Despite evident progress, less attention has been paid to their effectiveness in handling declarative paradigms such as Answer Set Programming (ASP), to date. In this paper we introduce BLAST: The first dedicated benchmarking methodology and associated dataset for evaluating the accuracy of LLMs in generating ASP code. BLAST provides a structured evaluation framework featuring two novel semantic metrics tailored to ASP code generation. The paper presents the results of an empirical evaluation involving ten well-established graph-related problems from the ASP literature and a diverse set of eight state-of-the-art LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。