用事件图结构建模世界,实现可解释的反事实推理。
Deterministic Event-Graph Substrates as World Models for Counterfactual Reasoning
- 将环境状态表示为可追加的三元组日志,通过结构化干预生成反事实分支。
- 在CLEVRER数据集上四项任务均优于符号基准,反事实任务领先0.80个百分点。
- 无需训练即可跨领域迁移,适合需要可解释性的决策系统研究。
我们研究事件图底座:一类将智能体状态表示为类型化的有向三元组追加日志的世界模型,通过结构化干预词汇表对反事实问题进行分叉处理。底座可在三元组层面进行检查,支持精确反事实推理,并在无学习组件的情况下跨领域迁移。我们形式化该类别,证明解释性与反事实查询之间的对偶性,两者均可归约为同一因果祖先遍历。评估了一个1,400行的CLEVRER-DSL解释器,基于通用底座运行于完整CLEVRER验证规模(n=75,618)。底座在四个问答类别中均超过NS-DR符号基准(分别领先9.89、20.26、17.65和0.80个百分点),在描述性和解释性任务上超越参数化基线ALOE,但在预测性和反事实任务上落后。此外,我们引入Twin-EventLog,一个包含500条规范的Park-canonical Smallville反事实基准,底座在全上下文条件下以18.80个百分点的优势超越Llama-3.1-8B。
原文摘要 · Abstract (English)
We study event-graph substrates: a class of world models that represent agent state as an append-only log of typed RDF triples and answer counterfactual queries by forking the log under a structured intervention vocabulary. Substrates are inspectable at the triple level, support exact counterfactuals, and transfer across domains without learned components. We formalize the class, prove a duality between explanatory and counterfactual queries that reduces both to the same causal-ancestor traversal, and evaluate a 1,400-line CLEVRER-DSL interpreter atop a domain-agnostic substrate runtime at full CLEVRER validation scale (n=75,618). The substrate exceeds the NS-DR symbolic oracle on all four per-question categories (by 9.89, 20.26, 17.65, and 0.80 percentage points), and exceeds the parametric ALOE baseline on descriptive and explanatory while lagging on predictive and counterfactual. We also introduce twin-EventLog, a 500-specification Park-canonical Smallville counterfactual benchmark on which the substrate exceeds Llama-3.1-8B with full context by 18.80 points joint accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。