呼吁构建完整无误的正式推理评估基准
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
- 主张建立开放代码、数据与完整无错的评测基准
- 指出当前评测实践易误导研究方向
- 适合自动化定理证明与形式化领域的研究者
本文针对形式推理与自动定理证明领域的评测实践展开批判性但建设性的讨论。我们主张,开放代码、开放数据以及完整且无错误的评测基准将加速该领域进展。文中识别出阻碍贡献的若干实践,并提出改进方案。同时,探讨了可能产生误导性评价信息的做法。旨在促进自动化定理证明、自动形式化与非形式推理相关研究者的协同交流。
原文摘要 · Abstract (English)
This position paper provides a critical but constructive discussion of current practices in benchmarking and evaluative practices in the field of formal reasoning and automated theorem proving. We take the position that open code, open data, and benchmarks that are complete and error-free will accelerate progress in this field. We identify practices that create barriers to contributing to this field and suggest ways to remove them. We also discuss some of the practices that might produce misleading evaluative information. We aim to create discussions that bring together people from various groups contributing to automated theorem proving, autoformalization, and informal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。