通过二进制内在维度检测大数据中的语义关联,突破维度诅咒限制。
Unsupervised detection of semantic correlations in big data
- 用二进制内在维度衡量数据语义复杂度,捕捉高维特征间隐含关联。
- 在磁性模型中成功识别相变点,验证了方法对关键结构的敏感性。
- 适用于图像与文本的深层神经网络语义分析,适合大尺度数据研究者。
真实世界数据以极高的特征向量形式存储,这些变量因多特征协同作用而存在复杂相关性,其模式对应语义角色,可被人类大脑与人工神经网络自然识别,从而实现基于上下文的图像或文本缺失部分预测。本文提出一种在二进制表示的高维数据中检测此类相关性的方法。通过估计数据集的二进制内在维度(binary intrinsic dimension),量化描述数据所需的最小独立坐标数,以此反映语义复杂度。该算法对维度诅咒不敏感,适用于大数据分析。我们测试其在模型磁性系统中识别相变的能力,并应用于深度神经网络内图像与文本的语义相关性检测。
原文摘要 · Abstract (English)
In real-world data, information is stored in extremely large feature vectors. These variables are typically correlated due to complex interactions involving many features simultaneously. Such correlations qualitatively correspond to semantic roles and are naturally recognized by both the human brain and artificial neural networks. This recognition enables, for instance, the prediction of missing parts of an image or text based on their context. We present a method to detect these correlations in high-dimensional data represented as binary numbers. We estimate the binary intrinsic dimension of a dataset, which quantifies the minimum number of independent coordinates needed to describe the data, and is therefore a proxy of semantic complexity. The proposed algorithm is largely insensitive to the so-called curse of dimensionality, and can therefore be used in big data analysis. We test this approach identifying phase transitions in model magnetic systems and we then apply it to the detection of semantic correlations of images and text inside deep neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。