评测8个团队的神经网络验证工具,推动验证技术标准化。
The 6th International Verification of Neural Networks Competition (VNN-COMP 2025): Summary and Results
- 采用ONNX和VNN-LIB标准格式统一模型与规范
- 在相同硬件上自动评测,8个团队参与16+9个基准测试
- 适合关注神经网络验证工具评估的开发者与研究者
本文总结了2025年第六届国际神经网络验证竞赛(VNN-COMP 2025)的成果。该竞赛作为第8届人工智能验证国际研讨会(SAIV)的一部分,与第37届计算机辅助验证国际会议(CAV)同期举行。竞赛旨在公平客观地对比前沿神经网络验证工具,推动工具接口标准化,并促进验证社区交流。为此,定义了统一的模型格式(ONNX)与规范格式(VNN-LIB),所有工具在等成本硬件上通过基于AWS实例的自动评测流程进行评估,参赛者需在最终测试集公开前提交参数配置。2025年共有8支队伍参与,覆盖16个常规基准和9个扩展基准。本报告详述了竞赛规则、基准集、参评工具、评测结果及经验教训。
原文摘要 · Abstract (English)
This report summarizes the 6th International Verification of Neural Networks Competition (VNN-COMP 2025), held as a part of the 8th International Symposium on AI Verification (SAIV), that was collocated with the 37th International Conference on Computer-Aided Verification (CAV). VNN-COMP is held annually to facilitate the fair and objective comparison of state-of-the-art neural network verification tools, encourage the standardization of tool interfaces, and bring together the neural network verification community. To this end, standardized formats for networks (ONNX) and specification (VNN-LIB) were defined, tools were evaluated on equal-cost hardware (using an automatic evaluation pipeline based on AWS instances), and tool parameters were chosen by the participants before the final test sets were made public. In the 2025 iteration, 8 teams participated on a diverse set of 16 regular and 9 extended benchmarks. This report summarizes the rules, benchmarks, participating tools, results, and lessons learned from this iteration of this competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。