对比人类与模型对图像的感知差异,发现二者各有优势且不可替代。
Perception of Visual Content: Differences Between Humans and Foundation Models
- 用多模态模型生成标注,与人类标注对比分析跨地区感知差异。
- 模型在区域分类和收入预测上表现更优,但人类在非动作类任务中更准。
- 研究揭示了人类标注的独特价值,提示其仍不可被完全替代。
人类标注常用于训练机器学习模型,但近年来语言和多模态基础模型被用来替代并扩展人工标注。本研究探讨不同社会经济背景下的图像标注中,人类与机器生成标注的相似性(RQ1),以及其对模型性能与偏见的影响(RQ2)。数据集包含来自不同地理区域和收入水平人群的日常活动与家庭环境图像。结果显示:在低层级特征(词汇类型与句式结构)上,人类与模型标注最相似;但在整体感知上,两者在各地区保持一致。模型生成的标注在区域分类任务中表现最佳,在收入回归任务中,模型物体与文本标注表现最优;在动作类别识别上,模型更优,而人类在非动作类别上更具优势。研究强调人类与机器标注均具重要价值,人类标注仍无法被取代。
原文摘要 · Abstract (English)
Human-annotated content is often used to train machine learning (ML) models. However, recently, language and multi-modal foundational models have been used to replace and scale-up human annotator's efforts. This study explores the similarity between human-generated and ML-generated annotations of images across diverse socio-economic contexts (RQ1) and their impact on ML model performance and bias (RQ2). We aim to understand differences in perception and identify potential biases in content interpretation. Our dataset comprises images of people from various geographical regions and income levels, covering various daily activities and home environments. ML captions and human labels show highest similarity at a low-level, i.e., types of words that appear and sentence structures, but all annotations are consistent in how they perceive images across regions. ML Captions resulted in best overall region classification performance, while ML Objects and ML Captions performed best overall for income regression. ML annotations worked best for action categories, while human input was more effective for non-action categories. These findings highlight the notion that both human and machine annotations are important, and that human-generated annotations are yet to be replaceable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。