提出可机器学习的离散集合新定义,让模型能高效识别与生成特定二进制串。
Machine-learnable Sets

- 用有界复杂度布尔自编码器定义集合的可学习性
- 在镜像翻转的罗夏图案等集合上验证了学习能力
- 提出迭代方法使复杂集合变为可学习,适合算法与认知研究者
本研究提出一类大型离散集合的形式化定义,其元素具备三个特征:易于识别、易于生成,且这些任务可通过少量样本高效学习。该形式化聚焦于二进制字符串集合,基于存在有界复杂度布尔自编码器来定义“机器可学习性”,该编码器能固定集合元素。实验中,自编码器由布尔阈值网络实现。在罗夏图案(镜像半边可能反色)等集合上验证了机器可学习性;对于更复杂的集合,其元素仅被允许自编码器近似固定。我们进一步提出一种简单迭代机制,可将这类“野生”集合逐步演化为真正可学习的集合。
原文摘要 · Abstract (English)
In this study we present a formal definition of large discrete sets having, informally, three properties: their elements are easily recognized, easily generated, and the latter tasks are easily learned from examples. The formalism is specialized to sets of binary strings and a definition of "machine-learnability" based on the existence of a bounded-complexity Boolean autoencoder that fixes the elements of the set. We present experiments where the autoencoders are implemented by nets of Boolean threshold functions. Machine-learnability is demonstrated for Rorschach patterns (that may have reversed contrast in the mirrored half), and considerably "wilder" sets whose elements are only approximately fixed by admissible autoencoders. In the second case we demonstrate a simple iteration that evolves wild sets to make them properly machine-learnable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。