测试视觉模型能否分离字母身份与位置,发现当前顶尖模型仍不及人类。
Disentanglement and Compositionality of Letter Identity and Letter Position in Variational Auto-Encoder Vision Models
- 用新基准CompOrth评估VAE模型对字母身份和位置的分离能力。
- 模型能区分图像中字母的上下左右位置,但无法识别字母在词中的实际位置。
- 揭示了现有模型在组合性推理和零样本泛化上的严重缺陷,适合研究神经网络认知局限的人看。
人类读者能准确数出单词中字母个数(如“buffalo”有7个字母)、从指定位置删除或添加字母。这表明大脑已学会分离字母身份与位置信息。这种分离是人类具备无限组合新字符串能力的基础。现代深度神经网络是否也具备此能力?本文测试了在书写单词图像上训练的β-变分自编码器(β-VAE)能否分离字母身份与位置。我们构建了新基准CompOrth,用于评估视觉模型在正字法上的组合学习与零样本泛化能力。通过一系列复杂度递增的测试,发现模型虽能有效分离如水平、垂直等表面特征(即图像中的‘视网膜’位置),却几乎无法分离字母的实际位置与身份,且完全缺乏对词长的概念。结果表明,当前最先进的β-VAE模型在组合性理解上远逊于人类,并提出了新的挑战与评估工具。
原文摘要 · Abstract (English)
Human readers can accurately count how many letters are in a word (e.g., 7 in ``buffalo''), remove a letter from a given position (e.g., ``bufflo'') or add a new one. The human brain of readers must have therefore learned to disentangle information related to the position of a letter and its identity. Such disentanglement is necessary for the compositional, unbounded, ability of humans to create and parse new strings, with any combination of letters appearing in any positions. Do modern deep neural models also possess this crucial compositional ability? Here, we tested whether neural models that achieve state-of-the-art on disentanglement of features in visual input can also disentangle letter position and letter identity when trained on images of written words. Specifically, we trained beta variational autoencoder ($β$-VAE) to reconstruct images of letter strings and evaluated their disentanglement performance using CompOrth - a new benchmark that we created for studying compositional learning and zero-shot generalization in visual models for orthography. The benchmark suggests a set of tests, of increasing complexity, to evaluate the degree of disentanglement between orthographic features of written words in deep neural models. Using CompOrth, we conducted a set of experiments to analyze the generalization ability of these models, in particular, to unseen word length and to unseen combinations of letter identities and letter positions. We found that while models effectively disentangle surface features, such as horizontal and vertical `retinal' locations of words within an image, they dramatically fail to disentangle letter position and letter identity and lack any notion of word length. Together, this study demonstrates the shortcomings of state-of-the-art $β$-VAE models compared to humans and proposes a new challenge and a corresponding benchmark to evaluate neural models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。