用极少标签实现物体中心推理,让神经符号系统更智能。
Weakly Supervised Concept Learning for Object-centric Visual Reasoning

- 用槽位架构+变分自编码器,弱监督下自动提取可解释符号。
- 仅需1%标签就能发现复杂抽象规则,且跨域鲁棒性强。
- 适合做少样本、可解释推理任务的研究者和工程师。
神经符号系统旨在结合深度神经网络对原始传感器输入的处理能力与符号人工智能的少样本性能。两阶段方法将基于DNN的感知与基于规则的推理显式解耦,避免了端到端可微方法的优化与可解释性问题,但需要昂贵的感知输出标注。本文提出一种高效的弱监督方案,用于感知阶段的符号接地,支持物体中心推理中的逻辑归纳。该方法结合基于槽位的架构与变分自编码器(VAE),在潜在空间中通过概念引导实现人类可解释的符号化。生成的预测被转化为符号背景知识,供归纳逻辑编程(ILP)、决策树和贝叶斯网络等推理框架使用。在合成数据集和真实世界数据集上的广泛实证评估表明,该方法能在仅1%标签监督下发现复杂的抽象规则,并在显著领域偏移下仍保持鲁棒性。值得注意的是,在1%监督条件下,其领域泛化性能甚至优于当前主流基础模型基线。
原文摘要 · Abstract (English)
Neurosymbolic systems promise to combine deep neural network's (DNN) processing of raw sensor inputs with few-shot performance of symbolic artificial intelligence. Two-stage approaches explicitly decouple DNN based perception from subsequent rule based reasoning. This avoids optimization and interpretability issues of end to end differentiable approaches, but requires costly labels for the perception output. This paper introduces an efficient weak supervision scheme for the perception stage to ground its output symbols for logical induction in object-centric reasoning tasks. It combines a slot-based architecture for object-centricity with a Variational Autoencoder (VAE) for self-supervision, competing with concept guidance on latent dimensions for human interpretable grounding. The resulting predictions are translated into symbolic background knowledge for reasoning frameworks, such as Inductive Logic Programming (ILP), Decision Trees, and Bayesian Networks. Our extensive empirical evaluation on synthetic and real world datasets shows that our approach can discover complex, abstract rules for object centric reasoning whilst reducing supervision to as little as 1% of labels, and being robust even under substantial domain shift. Notably, at 1% supervision it even outperforms state of the art foundation model baselines in domain generalization
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。