不依赖模型假设,通过比较数据发现隐藏概念
Nonparametric Identification of Latent Concepts
- 用跨类比较实现无参数概念识别
- 多样数据下可准确还原隐藏概念结构
- 理论适用于复杂场景,适合研究概念学习的学者
人类通过比较多样观察来学习概念,从而以组合方式理解世界并实现外推。本文认为,这种基于比较的认知机制对机器学习同样关键,有助于从数据中恢复真实概念,并为概念学习提供正确性保障。尽管该领域已有显著的实证成功,但缺乏普遍的理论支持。我们提出一种理论框架,可在多种观察类别下识别具有多个类别的概念,无需假设具体概念类型、函数关系或参数化生成模型。即使全局条件不满足,仍可通过局部比较为尽可能多的概念提供替代保证,扩展了理论的适用范围。此外,类别与概念之间的隐藏结构也可被非参数地识别。我们在合成数据和真实场景中验证了理论结果。
原文摘要 · Abstract (English)
We are born with the ability to learn concepts by comparing diverse observations. This helps us to understand the new world in a compositional manner and facilitates extrapolation, as objects naturally consist of multiple concepts. In this work, we argue that the cognitive mechanism of comparison, fundamental to human learning, is also vital for machines to recover true concepts underlying the data. This offers correctness guarantees for the field of concept learning, which, despite its impressive empirical successes, still lacks general theoretical support. Specifically, we aim to develop a theoretical framework for the identifiability of concepts with multiple classes of observations. We show that with sufficient diversity across classes, hidden concepts can be identified without assuming specific concept types, functional relations, or parametric generative models. Interestingly, even when conditions are not globally satisfied, we can still provide alternative guarantees for as many concepts as possible based on local comparisons, thereby extending the applicability of our theory to more flexible scenarios. Moreover, the hidden structure between classes and concepts can also be identified nonparametrically. We validate our theoretical results in both synthetic and real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。