arXiv:2512.01822cs.CLcs.AI2025-12被引 4

首个评估AI智能体创新潜力的基准,关注解法多样性而非仅正确性。

InnoGym: Benchmarking the Innovation Potential of AI Agents

  • 提出性能提升与方法新颖性双指标,衡量创新潜力。
  • 18个真实工程科学任务中,部分智能体方案新颖但稳定性差。
  • 适合研究AI创造力、智能体评估与具身智能的学者参考。

大语言模型和智能体在代码生成、数学推理和科学发现方面取得了显著进展,但现有基准主要衡量答案正确性,忽视了解法背后的多样性。真正的创新不仅在于结果正确,更在于方法的原创性。我们提出InnoGym,首个系统评估AI智能体创新潜力的基准与框架。InnoGym引入两个互补指标:性能提升(衡量相比最优已知解的改进程度)和新颖性(捕捉方法上与先前方案的差异)。基准包含18个来自真实工程与科学领域的精心设计任务,通过资源筛选、评估器验证和解法收集实现标准化。此外,我们提供iGym,一个统一的执行环境,支持可复现的长周期评估。大量实验表明,尽管某些智能体能产生新颖方法,但其鲁棒性不足导致性能提升有限。这些结果凸显了创造力与有效性之间的关键差距,强调了同时评估两者的重要性。

原文摘要 · Abstract (English)

LLMs and Agents have achieved impressive progress in code generation, mathematical reasoning, and scientific discovery. However, existing benchmarks primarily measure correctness, overlooking the diversity of methods behind solutions. True innovation depends not only on producing correct answers but also on the originality of the approach. We present InnoGym, the first benchmark and framework designed to systematically evaluate the innovation potential of AI agents. InnoGym introduces two complementary metrics: performance gain, which measures improvement over the best-known solutions, and novelty, which captures methodological differences from prior approaches. The benchmark includes 18 carefully curated tasks from real-world engineering and scientific domains, each standardized through resource filtering, evaluator validation, and solution collection. In addition, we provide iGym, a unified execution environment for reproducible and long-horizon evaluations. Extensive experiments show that while some agents produce novel approaches, their lack of robustness limits performance gains. These results highlight a key gap between creativity and effectiveness, underscoring the need for benchmarks that evaluate both.

AI创新智能体评估基准测试方法新颖性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。