arXiv:2505.09859cs.CV2025-05被引 1

用结构化表示和类比映射,仅凭少量样本就能学会复杂的视觉概念。

Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction

  • 通过深度学习对少量样本的结构化表示进行类比映射,构建组合概念模型
  • 在少样本场景下表现接近人类,优于使用非结构化特征的基线模型
  • 揭示了关系相似性在分类中的关键作用,适合认知建模与少样本学习研究

人类能够从极少的例子中学会新视觉概念,这是认知能力的重要体现。传统类别学习模型将每个例子表示为无结构的特征向量,而组合概念学习被认为依赖于(1)对例子的结构化表示(如包含对象及其关系的有向图),以及(2)通过类比映射识别例子间的共享关系结构。本文提出概率模式归纳(PSI)模型,利用深度学习对极少数样本的结构化表示进行类比映射,形成称为“模式”的组合概念。PSI采用一种新式相似性度量,同时权衡对象级相似性和关系级相似性,并引入机制放大对分类重要的关系,类似于传统模型中的选择性注意。实验表明,PSI展现出类人学习性能,显著优于两个对照组:使用深层网络提取的非结构化特征的原型模型,以及使用较弱结构化表示的PSI变体。值得注意的是,其类人表现源于一种自适应策略——逐步提升关系相似性权重,强化区分类别的关系贡献。结果表明,结构化表示与类比映射是建模快速类人组合视觉概念学习的关键,也展示了深度学习在构建心理模型中的潜力。

原文摘要 · Abstract (English)

The ability to learn new visual concepts from limited examples is a hallmark of human cognition. While traditional category learning models represent each example as an unstructured feature vector, compositional concept learning is thought to depend on (1) structured representations of examples (e.g., directed graphs consisting of objects and their relations) and (2) the identification of shared relational structure across examples through analogical mapping. Here, we introduce Probabilistic Schema Induction (PSI), a prototype model that employs deep learning to perform analogical mapping over structured representations of only a handful of examples, forming a compositional concept called a schema. In doing so, PSI relies on a novel conception of similarity that weighs object-level similarity and relational similarity, as well as a mechanism for amplifying relations relevant to classification, analogous to selective attention parameters in traditional models. We show that PSI produces human-like learning performance and outperforms two controls: a prototype model that uses unstructured feature vectors extracted from a deep learning model, and a variant of PSI with weaker structured representations. Notably, we find that PSI's human-like performance is driven by an adaptive strategy that increases relational similarity over object-level similarity and upweights the contribution of relations that distinguish classes. These findings suggest that structured representations and analogical mapping are critical to modeling rapid human-like learning of compositional visual concepts, and demonstrate how deep learning can be leveraged to create psychological models.

少样本学习组合概念认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。