用大模型分析社交媒体图片中人与自然互动,构建新数据集并评估多种方法效果。
Exploring Social Media Image Categorization Using Large Models with Different Adaptation Methods: A Case Study on Cultural Nature's Contributions to People
- 构建FLIPS数据集,聚焦人类与自然互动的社交媒体图像。
- 对比不同大模型组合在成本、效率和准确率上的表现。
- 为文化遗产与公众互动研究提供可复现的评估框架,适合人文地理与AI交叉研究者。
社交媒体图像为建模和理解人类与自然及文化遗产的互动提供了宝贵视角。然而,由于视觉内容多样且开放世界特征明显,对这些图像进行语义分类仍极具挑战性,尤其当类别涉及抽象概念且缺乏一致视觉模式时。现有研究多依赖人工标注,且缺乏公开基准数据集,难以进行有效比较。随着大语言模型(LLMs)、大视觉模型(LVMs)和大视觉语言模型(LVLMs)的持续发展,为解决该问题提供了广阔的新路径。本文提出:1)构建一个名为FLIPS的Flickr图像数据集,涵盖人类与自然互动场景;2)评估基于不同类型及组合的大模型,采用多种适配方法的解决方案。从成本、生产效率、可扩展性和结果质量等方面系统评估其性能,以应对社交媒体图像分类的挑战。
原文摘要 · Abstract (English)
Social media images provide valuable insights for modeling, mapping, and understanding human interactions with natural and cultural heritage. However, categorizing these images into semantically meaningful groups remains highly complex due to the vast diversity and heterogeneity of their visual content as they contain an open-world human and nature elements. This challenge becomes greater when categories involve abstract concepts and lack consistent visual patterns. Related studies involve human supervision in the categorization process and the lack of public benchmark datasets make comparisons between these works unfeasible. On the other hand, the continuous advances in large models, including Large Language Models (LLMs), Large Visual Models (LVMs), and Large Visual Language Models (LVLMs), provide a large space of unexplored solutions. In this work 1) we introduce FLIPS a dataset of Flickr images that capture the interaction between human and nature, and 2) evaluate various solutions based on different types and combinations of large models using various adaptation methods. We assess and report their performance in terms of cost, productivity, scalability, and result quality to address the challenges of social media image categorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。