arXiv:2606.02556cs.CL2026-06

用文本游戏测试智能体规则推理能力,发现大模型在复杂任务上仍存瓶颈。

HERO'S JOURNEY: Testing Complex Rule Induction with Text Games

论文配图:HERO'S JOURNEY: Testing Complex Rule Induction with Text Games
图 1 · 摘自论文原文
  • 设计八类任务,通过多步执行检验隐藏规则推理能力
  • 大模型能部分推断规则,但执行环节成主要瓶颈
  • 属性类任务可提升,程序类任务仍难突破,适合研究规则学习者

我们提出HERO'S JOURNEY,一个面向目标导向的叙事性任务中规则归纳的基准测试。该基准包含八个任务,涵盖属性与过程归纳两类,每类有四种结构化规则形式,支持可控词汇语义和可辨识性条件。评估当前主流大语言模型发现,模型具备一定规则归纳能力,但表现有限且任务间差异明显;过程执行成为主要瓶颈,而表面语义影响较小。针对归纳的定向调控方法在属性任务上有效,但在程序任务上无稳定提升,表明程序归纳仍是开放挑战。

原文摘要 · Abstract (English)

We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations and act on them through multi-step execution. HERO'S JOURNEY covers eight tasks across attribute and procedural induction families, each with four structural rule forms, controllable lexical grounding, and identifiability conditions. Evaluating state-of-the-art LLMs, we find that models show evidence of rule induction, but the ability is limited and uneven across tasks. Meanwhile, process execution adds an execution bottleneck for models, whereas surface semantics has minimal effect. Induction-specific steering methods improve performance on attribute tasks but show no reliable gains on procedural tasks, suggesting the gap in procedural induction remains an open challenge.

规则推理大模型评估文本游戏任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。