用大模型自动辅助安全研究可复现性评估,提升评审效率。
Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers
- 用大模型自动判断论文代码可复现性并生成运行环境
- 可复现性评分准确率超72%,28%的代码能自动部署
- 精准识别7类常见方法缺陷,适合会议评审使用
可复现性评估(AE)对保障研究透明性和可靠性至关重要,尤其在物联网与网络物理系统等安全敏感领域,面对大规模异构数据和关键控制行为,更需确保研究结果真实可信。然而传统人工复现检查耗时且难以应对日益增长的投稿量。本文证明大语言模型可在三项任务中提供有力支持:(i)基于文本的可复现性评分,(ii)自主构建沙箱运行环境,(iii)识别方法论缺陷。实验显示,该评估工具可实现超过72%的准确率,能为28%的可运行安全研究资源自动搭建执行环境,且对七类常见问题的检测F1值均高于92%。该工具显著降低审稿人工作量,若融入现有评审流程,有望激励作者提交更高质、更可复现的研究成果。物联网、嵌入式系统及网络安全领域的会议可将其纳入评审体系,用于决定是否授予可复现性徽章,从而提升整个学术生态的可持续性。
原文摘要 · Abstract (English)
Artifact Evaluation (AE) is essential for ensuring the transparency and reliability of research, closing the gap between exploratory work and real-world deployment is particularly important in cybersecurity, particularly in IoT and CPSs, where large-scale, heterogeneous, and privacy-sensitive data meet safety-critical actuation. Yet, manual reproducibility checks are time-consuming and do not scale with growing submission volumes. In this work, we demonstrate that Large Language Models (LLMs) can provide powerful support for AE tasks: (i) text-based reproducibility rating, (ii) autonomous sandboxed execution environment preparation, and (iii) assessment of methodological pitfalls. Our reproducibility-assessment toolkit yields an accuracy of over 72% and autonomously sets up execution environments for 28% of runnable cybersecurity artifacts. Our automated pitfall assessment detects seven prevalent pitfalls with high accuracy ($F_1$ > 92%). Hence, the toolkit significantly reduces reviewer effort and, when integrated into established AE processes, could incentivize authors to submit higher-quality and more reproducible artifacts. IoT, CPS, and cybersecurity conferences and workshops may integrate the toolkit into their peer-review processes to support reviewers' decisions on awarding artifact badges, improving the overall sustainability of the process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。