用符号推理指导神经模型,更好回答视频中的假设性问题。
Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for Counterfactual Question Answering
- 构建因果图,用逻辑编程协调感知与模拟模块。
- 在CLEVRER上达当前最佳,显著超越已有模型。
- 适合需要严谨因果推断的视频理解研究者。
关于视频动态的因果与时间推理是一项挑战性任务。尽管结合符号推理与神经感知预测的神经符号模型展现出潜力,但在回答反事实问题时仍存在局限。本文提出一种增强神经符号模型的方法,通过事件间的因果关系进行符号推理。定义因果图表示这些关系,并采用答案集编程(ASP)这一声明式逻辑编程方法,确定如何协调感知与仿真模块。在CLEVRER和CRAFT两个基准上验证了该方法的有效性。在CLEVRER上达到当前最优性能,显著优于现有模型。在CRAFT上,利用GPT-3.5和GPT-4等大语言模型作为动力学模拟器的代理,通过符号因果推理生成引导提示,进一步提升反事实问题的回答效果。
原文摘要 · Abstract (English)
Causal and temporal reasoning about video dynamics is a challenging problem. While neuro-symbolic models that combine symbolic reasoning with neural-based perception and prediction have shown promise, they exhibit limitations, especially in answering counterfactual questions. This paper introduces a method to enhance a neuro-symbolic model for counterfactual reasoning, leveraging symbolic reasoning about causal relations among events. We define the notion of a causal graph to represent such relations and use Answer Set Programming (ASP), a declarative logic programming method, to find how to coordinate perception and simulation modules. We validate the effectiveness of our approach on two benchmarks, CLEVRER and CRAFT. Our enhancement achieves state-of-the-art performance on the CLEVRER challenge, significantly outperforming existing models. In the case of the CRAFT benchmark, we leverage a large pre-trained language model, such as GPT-3.5 and GPT-4, as a proxy for a dynamics simulator. Our findings show that this method can further improve its performance on counterfactual questions by providing alternative prompts instructed by symbolic causal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。