用Jupyter笔记本当基准,测大模型生成代码能力
Themisto: Jupyter-Based Runtime Benchmark
- 构建包含运行时信息的Jupyter笔记本基准数据集
- 现有大模型在代码预测和生成任务上表现不佳
- 强调运行时上下文对代码模型开发的重要性
本文提出一个基于Jupyter笔记本开发轨迹的基准,用于衡量大语言模型(LLMs)利用运行时信息进行代码输出预测和代码生成的能力。我们证明当前一代的LLMs在这类任务中表现较差,并指出代码模型开发中一个显著被忽视的领域——即如何融入运行时上下文。该基准旨在推动对动态执行环境与代码生成之间关系的研究。
原文摘要 · Abstract (English)
In this work, we present a benchmark that consists of Jupyter notebooks development trajectories and allows measuring how large language models (LLMs) can leverage runtime information for predicting code output and code generation. We demonstrate that the current generation of LLMs performs poorly on these tasks and argue that there exists a significantly understudied domain in the development of code-based models, which involves incorporating the runtime context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。