构建D&D战斗战术推理基准,测试复杂规则下的决策能力
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

- 基于2014版规则文档构建战斗模拟基准,保留真实规则交互
- 分遭遇战与全天追踪两赛道,评估即时与长期资源管理能力
- 适合评估大模型在复杂规则下的战术推理与资源规划能力
游戏和模拟器通过将决策转化为可衡量结果,成为有价值的基准,但多数现有基准未能充分测试规则密集型战术推理——即在几何、时机、资源、目标及规则交互同时作用下做出优质选择的能力。我们提出DungeonBench,一个面向《龙与地下城》战斗的战术推理基准,覆盖绝大多数2014版系统参考文档中可由模拟器处理的战斗相关内容,保留简化模拟器常忽略的机制。每一步暴露完整的战术观察、待决策项及包含移动、攻击、法术、反应、目标、准备和稀缺资源的可执行选项列表。任务是评估合法选择的价值,其后果取决于行动经济、生物特性、战场几何、时间窗口及未来遭遇。该基准含两个赛道:遭遇战,评估单场战斗中的局部战术;全天,通过持续生命值、法术位、消耗品、准备状态和短休时间链接遭遇,迫使策略在即时优势与未来生存间权衡。同一引擎生成的决策流支持启发式控制器、语言模型策略、学习选项排序器和掩码动作强化学习代理。我们在共享决策流上评估前沿语言模型策略,结果表明完整战术观察未使基准饱和:前沿模型常赢直接遭遇,但在连环遭遇日中暴露出资源预算、休息时机和规则感知战术纪律的缺陷。
原文摘要 · Abstract (English)
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At each step, DungeonBench exposes a complete tactical observation, a pending decision, and an indexed list of executable options spanning movement, attacks, spells, reactions, objectives, preparation, and scarce resources. The task is to value legal choices whose consequences depend on action economy, creature traits, battlefield geometry, timing windows, and future encounters. DungeonBench has two tracks: Encounter, which evaluates local tactical play in single fights, and Day, which links encounters through persistent hit points, spell slots, consumables, preparation, and short-rest timing, forcing policies to trade off immediate tactical advantage against future survivability. The same engine-generated decision stream supports heuristic controllers, language-model policies, learned option rankers, and masked-action reinforcement-learning agents. We evaluate frontier language-model policies on this shared decision stream. Results show that full tactical observations do not saturate the benchmark: frontier policies often win direct encounters, but linked encounter days expose failures in resource budgeting, rest timing, and rule-aware tactical discipline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。