构建游戏化基准,测试AI发现未知规律的能力
DiG-bench: Discovery in Games
- 设计70个需实验探索规则的游戏,通过互动发现隐藏机制
- 最高难度挑战顶尖智能体,人类可首试通关全部游戏
- 适合研究自主探索与科学发现的AI系统,支持安全评估
发现——提出新的普遍规律——是科学过程的核心。尽管重要,当前人工智能评估体系中缺乏直接检验在受控环境中通过实验发现新知识能力的基准,且目标未知。为此,我们发布新基准DiG-bench(Discovery in Games)。该基准包含70个独立游戏,每个以短字符串编码,拥有独特变换规则,需通过交互与实验揭示。关卡设置一系列挑战,用于测试规则是否被掌握,且每关获胜条件未知。游戏按七级难度分级,最低级可被多个模型常规解决,最高级则挑战最先进智能体系统。所有70个游戏均至少被一人类在首次尝试中成功通关。其中21个公开发布,其余保留私有以保障评估安全性。
原文摘要 · Abstract (English)
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。