arXiv:2508.00869cs.LGcs.ET2025-08

用稀疏二进制向量构建离散编码空间,模拟语言与医学图像的结构特征。

Discrete approach to machine learning

  • 用固定长度线性向量实现低复杂度的随机降维。
  • 在俄语、英语和免疫组化数据上验证了编码空间的几何结构。
  • 发现编码布局与哺乳动物新皮层的螺旋状结构相似,或具生物启发意义。

本文提出一种基于稀疏比特向量和固定长度线性向量的编码与结构信息处理方法。展示了针对多维代码空间与线性空间的离散推测性随机降维方法,具有线性渐近复杂度;并提出一种几何方法,用于获取反映特定模态内部结构的离散嵌入。通过形态学(俄语、英语)及免疫组化标记三种模态实例,研究了代码空间的结构与特性。结果发现代码空间布局与哺乳动物新皮层中的‘螺旋状’结构存在相似性,作者谨慎推测该现象可能暗示新皮层组织与模型中过程间的潜在共性。

原文摘要 · Abstract (English)

The article explores an encoding and structural information processing approach using sparse bit vectors and fixed-length linear vectors. The following are presented: a discrete method of speculative stochastic dimensionality reduction of multidimensional code and linear spaces with linear asymptotic complexity; a geometric method for obtaining discrete embeddings of an organised code space that reflect the internal structure of a given modality. The structure and properties of a code space are investigated using three modalities as examples: morphology of Russian and English languages, and immunohistochemical markers. Parallels are drawn between the resulting map of the code space layout and so-called pinwheels appearing on the mammalian neocortex. A cautious assumption is made about similarities between neocortex organisation and processes happening in our models.

离散表示编码空间神经启发语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。