探究视觉Transformer与人类感知的对齐程度,发现模型越大越不贴近人眼判断。
Do Vision Transformers See Like Humans? Evaluating their Perceptual Alignment
- 通过TID2013数据集评估模型规模、训练策略对感知对齐的影响。
- 模型越大、训练重复次数越多,与人类判断的偏差越大。
- 数据增强和正则化会进一步削弱感知对齐,适合关注人机一致性场景。
视觉Transformer(ViTs)在图像识别任务中表现优异,但其与人类感知的对齐程度尚未深入探索。本研究系统分析了模型规模、数据集规模、数据增强和正则化对ViT在TID2013数据集上感知对齐的影响。结果表明,更大模型表现出更低的感知对齐度,与已有研究一致。增加数据集多样性影响较小,但重复使用相同图像进行训练会降低对齐度。更强的数据增强和正则化也进一步削弱对齐,尤其在多次训练循环的模型中更为明显。这些发现揭示了模型复杂性、训练策略与人类感知对齐之间的权衡,对需要类人视觉理解的应用具有重要启示。
原文摘要 · Abstract (English)
Vision Transformers (ViTs) achieve remarkable performance in image recognition tasks, yet their alignment with human perception remains largely unexplored. This study systematically analyzes how model size, dataset size, data augmentation and regularization impact ViT perceptual alignment with human judgments on the TID2013 dataset. Our findings confirm that larger models exhibit lower perceptual alignment, consistent with previous works. Increasing dataset diversity has a minimal impact, but exposing models to the same images more times reduces alignment. Stronger data augmentation and regularization further decrease alignment, especially in models exposed to repeated training cycles. These results highlight a trade-off between model complexity, training strategies, and alignment with human perception, raising important considerations for applications requiring human-like visual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。