arXiv:2409.06509cs.CVcs.AI2024-09被引 3

让机器视觉模型更像人脑,从细粒度到粗粒度层次统一认知。

Aligning Machine and Human Visual Representations across Abstraction Levels

  • 用人类判断训练教师模型,再将人类认知结构迁移到视觉模型中
  • 在多层级语义抽象任务上,模型行为与人类更一致,不确定性预测更准
  • 提升模型泛化能力,适合追求可解释性与鲁棒性的AI系统研究者

深度神经网络在诸多视觉任务中表现优异,但其训练方式与人类学习存在根本差异,且泛化能力常不及人类。本文指出关键差距:人类概念知识按从细到粗的层次组织,而模型表征未能准确捕捉所有抽象层级。为此,先训练教师模型模仿人类判断,再通过微调将人类对齐的结构注入预训练视觉基础模型。经优化的模型在跨多层级语义抽象的相似性任务上更贴近人类行为与不确定性判断,且在多样机器学习任务中提升泛化与分布外鲁棒性。结果表明,融合人类知识使模型兼具人类认知一致性与实际应用价值,为构建更稳健、可解释、人性化的智能系统铺路。

原文摘要 · Abstract (English)

Deep neural networks have achieved success across a wide range of applications, including as models of human behavior and neural representations in vision tasks. However, neural network training and human learning differ in fundamental ways, and neural networks often fail to generalize as robustly as humans do raising questions regarding the similarity of their underlying representations. What is missing for modern learning systems to exhibit more human-aligned behavior? We highlight a key misalignment between vision models and humans: whereas human conceptual knowledge is hierarchically organized from fine- to coarse-scale distinctions, model representations do not accurately capture all these levels of abstraction. To address this misalignment, we first train a teacher model to imitate human judgments, then transfer human-aligned structure from its representations to refine the representations of pretrained state-of-the-art vision foundation models via finetuning. These human-aligned models more accurately approximate human behavior and uncertainty across a wide range of similarity tasks, including a new dataset of human judgments spanning multiple levels of semantic abstractions. They also perform better on a diverse set of machine learning tasks, increasing generalization and out-of-distribution robustness. Thus, infusing neural networks with additional human knowledge yields a best-of-both-worlds representation that is both more consistent with human cognitive judgments and more practically useful, thus paving the way toward more robust, interpretable, and human-aligned artificial intelligence systems.

视觉表征人类对齐泛化能力认知结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。