arXiv:2412.19985cs.LGcs.AI2024-12被引 72

8支团队比拼神经网络验证工具,12个标准+8个扩展基准测试。

The Fifth International Verification of Neural Networks Competition (VNN-COMP 2024): Summary and Results

  • 采用ONNX与VNN-LIB统一格式,确保公平评测。
  • 8支团队在同等硬件上完成12个常规+8个扩展测试。
  • 适合关注神经网络验证技术的科研与工程人员。

本文总结了第五届国际神经网络验证竞赛(VNN-COMP 2024),该竞赛作为第七届人工智能验证国际研讨会(SAIV)的一部分,与第36届计算机辅助验证国际会议(CAV)同期举办。赛事旨在推动先进神经网络验证工具的公平、客观比较,促进工具接口标准化,并凝聚验证领域研究力量。为此,竞赛定义了统一的模型格式(ONNX)与规范格式(VNN-LIB),所有工具在等成本硬件上通过基于AWS实例的自动评估流程进行测试,且参赛者需提前提交工具参数。2024年共有8支团队参与,涵盖12个常规基准和8个扩展基准。本报告总结了竞赛规则、基准集、参评工具、测试结果及经验教训。

原文摘要 · Abstract (English)

This report summarizes the 5th International Verification of Neural Networks Competition (VNN-COMP 2024), held as a part of the 7th International Symposium on AI Verification (SAIV), that was collocated with the 36th International Conference on Computer-Aided Verification (CAV). VNN-COMP is held annually to facilitate the fair and objective comparison of state-of-the-art neural network verification tools, encourage the standardization of tool interfaces, and bring together the neural network verification community. To this end, standardized formats for networks (ONNX) and specification (VNN-LIB) were defined, tools were evaluated on equal-cost hardware (using an automatic evaluation pipeline based on AWS instances), and tool parameters were chosen by the participants before the final test sets were made public. In the 2024 iteration, 8 teams participated on a diverse set of 12 regular and 8 extended benchmarks. This report summarizes the rules, benchmarks, participating tools, results, and lessons learned from this iteration of this competition.

神经网络验证竞赛评测工具对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。