用元学习模拟人类视觉的灵活适应能力,让模型像人一样从少量样本学会新任务。
Meta-learning as a principle for human-like visual representations

- 通过元学习训练模型,使其能从少量样本中快速掌握新任务。
- 元学习模型在预测人类相似性判断、语义规则学习上优于传统预训练模型。
- 该方法揭示了人类视觉系统灵活性的内在机制,适合研究认知科学与通用视觉模型者参考。
人类视觉表征的结构支撑了我们适应性行为的能力。尽管预训练神经网络在建模人类视觉表征上取得了前所未有的成功,但仍存在显著差距。我们提出一个原因:这些网络优化单一固定目标,而人类表征需支持开放性任务。我们假设这种灵活性源于元学习(学习如何学习),即一种推动表征从少量观察中习得新任务的压力。为验证此假设,我们在数千个语义丰富的任务上训练了一个序列模型,这些任务将图像映射到高层次概念,且未使用任何人类数据监督。相比其预训练基础编码器,元学习得到的表征在预测人类相似性判断、语义规则学习及高层视觉皮层活动方面表现更优。行为性能提升依赖于解耦的高层任务分布,而大脑对齐主要由学习如何学习的压力驱动。结果表明,人类视觉表征的灵活性反映了即时学习新语义关系的功能需求。
原文摘要 · Abstract (English)
The structure of human visual representations underpins our capacity for adaptive behaviour. While pretrained neural networks model human visual representations with unprecedented success, a large discrepancy remains. We propose one reason: these networks optimise a single fixed objective, whereas human representations must support open-ended tasks. We hypothesise this flexibility arises from meta-learning (learning to learn), a pressure shaping representations to acquire new tasks from few observations. To test this, we train a sequence model, without any supervision from human data, across thousands of semantically rich tasks mapping images to high-level concepts. Compared to their pretrained base encoders, meta-learned representations better predict human similarity judgements, semantic rule learning, and high-level visual cortex. Behavioural gains depend on disentangled, high-level task distributions, while brain alignment is driven primarily by the learning-to-learn pressure. Our results suggest the flexibility of human visual representations reflects the functional demand to learn new semantic relationships on the fly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。