用GPT-4o分析图像吸引力,发现其能部分模拟人类注意力机制。
Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
- 通过对比人类评分与GPT-4o预测,评估其对视觉吸引力的理解能力。
- 在图像对排序任务中,GPT-4o表现优于现有方法,可有效生成标注数据。
- 研究成果有助于构建更贴近人类兴趣的推荐与生成系统。
我们的日常生活深受所见内容的影响。吸引并维持注意力——即(视觉)有趣性的定义——至关重要。大规模多模态模型(LMMs)在海量视觉与文本数据上训练后展现出卓越能力。本文探索这些模型理解视觉有趣性概念的程度,并通过对比分析检验人类评估与GPT-4o(一种领先LMM)预测之间的对齐情况。研究发现,人类与GPT-4o之间存在部分对齐,且其表现优于现有最先进方法。因此,可用于高效标注图像对的有趣性,作为训练数据,将知识蒸馏至一个学习排序模型中。该研究为深入理解人类兴趣提供了新路径。
原文摘要 · Abstract (English)
Our daily life is highly influenced by what we consume and see. Attracting and holding one's attention -- the definition of (visual) interestingness -- is essential. The rise of Large Multimodal Models (LMMs) trained on large-scale visual and textual data has demonstrated impressive capabilities. We explore these models' potential to understand to what extent the concepts of visual interestingness are captured and examine the alignment between human assessments and GPT-4o's, a leading LMM, predictions through comparative analysis. Our studies reveal partial alignment between humans and GPT-4o. It already captures the concept as best compared to state-of-the-art methods. Hence, this allows for the effective labeling of image pairs according to their (commonly) interestingness, which are used as training data to distill the knowledge into a learning-to-rank model. The insights pave the way for a deeper understanding of human interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。