发现神经网络隐空间凸性与人类认知对齐有关
Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks
- 用人类行为数据检验模型表征的凸性与认知对齐关系
- 预训练和微调模型中凸性与对齐度呈正相关
- 适合关注可解释性与人机协同的AI研究者
理解神经网络如何与人类认知过程对齐,是构建更可解释、更可靠AI系统的关键一步。基于人类认知理论,本研究通过行为数据考察了预训练和微调视觉变换器模型中神经网络表征的凸性与人类-机器对齐之间的关系。研究发现,神经网络隐空间中形成的凸区域在一定程度上与人类定义的类别一致,并反映了人类在认知任务中使用的相似性关系。虽然优化对齐通常会增强凸性,但通过微调提升凸性对对齐的影响并不一致,表明二者关系复杂。本研究为理解隐空间凸性与人类-机器对齐的关系迈出了第一步。
原文摘要 · Abstract (English)
Understanding how neural networks align with human cognitive processes is a crucial step toward developing more interpretable and reliable AI systems. Motivated by theories of human cognition, this study examines the relationship between \emph{convexity} in neural network representations and \emph{human-machine alignment} based on behavioral data. We identify a correlation between these two dimensions in pretrained and fine-tuned vision transformer models. Our findings suggest that the convex regions formed in latent spaces of neural networks to some extent align with human-defined categories and reflect the similarity relations humans use in cognitive tasks. While optimizing for alignment generally enhances convexity, increasing convexity through fine-tuning yields inconsistent effects on alignment, which suggests a complex relationship between the two. This study presents a first step toward understanding the relationship between the convexity of latent representations and human-machine alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。