用极粗粒度标签也能训练出贴近人类视觉的AI模型
An extremely coarse feedback signal is sufficient for learning human-aligned visual representations
- 用主成分分析分割图像为2到64类,测试不同粗细标签对模型的影响
- 仅需8类粗标签训练的模型,其神经表征与人脑匹配度超过千类细标签模型
- 粗标签模型更接近人类感知相似性,适合追求人机对齐的AI设计
在视觉任务上训练的人工神经网络会发展出类似灵长类视觉系统的内部表征,这一发现指导了十年的计算神经科学。现有研究逐步采用更细粒度的监督信号,从物体分类到对比自监督目标,但监督信号精细程度对大脑对齐的影响仍不清楚。本文系统研究学习信号粗细如何影响与人类视觉的表征对齐。通过基于PCA的预训练嵌入分割,将训练图像分为2、4、8、16、…、64个类别,构建不同粒度的分类任务。在卷积和Transformer架构上训练数百个神经网络,并将其表征与猕猴电生理记录和人类fMRI反应进行比较。结果表明,仅需区分8个宽泛类别,模型就能获得与人脑匹配度不低于1000类细粒度模型的表征。更令人惊讶的是,这些粗粒度训练模型比所有其他评估模型(包括细粒度监督或自监督模型及主流大模型)更贴近人类感知相似性判断。这表明人类类视觉表征可由极粗反馈信号生成,重新定义了视觉学习所需信号的形态,为构建更符合人类感知的AI系统开辟新路径。
原文摘要 · Abstract (English)
Artificial neural networks trained on visual tasks develop internal representations resembling those of the primate visual system, a discovery that has guided a decade of computational neuroscience. Research on building brain-aligned models has progressively embraced finer-grained supervisory signals, from object classification to contrastive self-supervised objectives that maximize distinctions among individual images, yet the role of supervisory signal granularity on brain alignment remains largely unexamined. Here we systematically investigate how the coarseness of a learning signal shapes representational alignment with human vision. We parametrically vary the level of signal granularity using a data-driven approach that partitions a set of training images into varied numbers of categories (2, 4, 8, 16, ..., 64) via PCA-based splits of pretrained embeddings. We train hundreds of neural networks across convolutional and transformer architectures on these coarse classification tasks and compare their representations to macaque electrophysiology recordings and human fMRI responses. We find that networks trained to distinguish as few as 8 broad categories learn representations that match or exceed the neural alignment of models distinguishing 1,000-classes. Even more strikingly, these coarsely trained networks align more closely with human perceptual similarity judgments than all other models evaluated, including networks trained with fine-grained supervision or self-supervision as well as leading large-scale vision models. These results demonstrate that human-like visual representations emerge from remarkably coarse feedback, reframing what learning signals vision may require and opening a path toward building AI systems that are more aligned with human perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。