神经网络处理图像时自发产生语言统计规律。
Spontaneous emergence of linguistic statistical laws in images via artificial neural networks
- 用预训练网络提取图像像素,生成类文本单位。
- 这些单位符合齐普夫、希普斯和本福德定律。
- 无需符号设计,连接主义网络即可自动生成结构化表示。
作为文化核心要素,图像将感知转化为结构化表征,并经历类似自然语言的演化。鉴于视觉输入占人类感官体验的60%,自然会提出图像是否遵循与语言系统类似的统计规律。基于符号锚定理论(符号源于感知),我们将图像视为以视觉为中心的产物,利用预训练神经网络建模视觉处理过程。通过检测卷积核激活并提取像素,我们获得类文本单元,发现这些图像衍生表征遵循齐普夫律、希普斯律和本福德律等统计规律,与语言数据相似。值得注意的是,这些统计规律是自发出现的,无需显式符号或混合架构。结果表明,连接主义网络仅通过感知处理即可自动发展出结构化的准符号单元,暗示文本与符号特性可自然从神经网络中涌现,为理解提供了新视角。
原文摘要 · Abstract (English)
As a core element of culture, images transform perception into structured representations and undergo evolution similar to natural languages. Given that visual input accounts for 60% of human sensory experience, it is natural to ask whether images follow statistical regularities similar to those in linguistic systems. Guided by symbol-grounding theory, which posits that meaningful symbols originate from perception, we treat images as vision-centric artifacts and employ pre-trained neural networks to model visual processing. By detecting kernel activations and extracting pixels, we obtain text-like units, which reveal that these image-derived representations adhere to statistical laws such as Zipf's, Heaps', and Benford's laws, analogous to linguistic data. Notably, these statistical regularities emerge spontaneously, without the need for explicit symbols or hybrid architectures. Our results indicate that connectionist networks can automatically develop structured, quasi-symbolic units through perceptual processing alone, suggesting that text- and symbol-like properties can naturally emerge from neural networks and providing a novel perspective for interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。