arXiv:2607.20402cs.AI2026-07

让视觉推理同时学习感知、知识和逻辑,全程可微分训练。

SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data

论文配图:SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data
图 1 · 摘自论文原文
  • 用软符号张量表示推理状态,实现感知与逻辑的可微衔接。
  • 在KVQA任务上达成92.3%准确率,支持端到端可训练的逻辑推导。
  • 适合需要融合视觉理解与知识推理的研究者使用。

许多推理问题中,前提并非以离散符号形式出现,而是需从高维输入中推断得出。同时,谓词词汇、参数结构及可信证据由知识图谱(KG)或规则定义提供。传统神经符号系统在感知与推理之间存在离散接口。本文提出一种神经-软符号推理架构(SoftReason),可在潜在感知事实与知识提供的谓词间实现可微分的演绎推理。SoftReason通过将推理状态表示为候选常量与谓词上的局部软解释张量,消除梯度断裂。感知模块生成概率性基础事实,KG三元组作为高置信度软证据输入,每个查询锚点、谓词选择与闭包更新均保持可微。核心创新在于学习了一个可微的即时结论算子:利用谓词定义嵌入与潜在组合通道,形成软体-谓词混合,对所有可能见证者聚合,生成条件查询头事实,并通过单调概率或操作更新解释。该框架在知识感知视觉问答(KVQA)任务上实现端到端感知定位、知识图谱证据注入与可微闭包推导,完全可训练。

原文摘要 · Abstract (English)

In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or rule definitions. Classical neuro-symbolic pipelines have a discrete interface between perception and deduction. We present a neuro-soft-symbolic architecture for differentiable deductive reasoning over latent perceptual facts and knowledge-provided predicates. SoftReason removes the gradient gap by representing the deductive state as a local soft interpretation tensor over candidate constants and predicates. Perception proposes probabilistic base facts, KG triples enter as high-confidence soft evidence, and every query anchor, predicate choice, and closure update remains differentiable. Our core innovation is a learned differentiable lift of the immediate-consequence operator. It uses predicate-definition embeddings and latent composition channels to form soft body-predicate mixtures, aggregate over all possible witnesses, propose query-conditioned head facts, and update the interpretation through a monotone probabilistic OR. We instantiate the framework on Knowledge-aware Visual Question Answering (KVQA), and demonstrates how SoftReason supports end-to-end perceptual grounding, KG evidence injection, and differentiable deductive closure in one trainable architecture.

神经符号可微推理视觉问答知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。