用自然语言自动发现图像分类标准并组织无序图库。
Organizing Unstructured Image Collections using Natural Language
- 以文本为推理代理,从图像中自发提出语义分组标准
- 在COCO-4C和Food-4C数据集上成功识别出4类有意义的聚类
- 适用于发现生成模型偏见与社交媒体图像传播规律
本文提出并研究了开放式的多语义聚类任务(OpenSMC)。给定一个大规模无结构的图像集合,目标是无需人工干预,自动发现多个多样化的语义聚类标准(如行为或地点),并据此对图像进行组织。我们提出的X-Cluster框架将文本作为推理代理:同时扫描整个图像集合,以自然语言形式提出候选分类标准,并按每种标准将图像划分为有意义的簇。这与以往假设预定义聚类标准或固定簇数的方法截然不同。为评估X-Cluster,我们构建了两个新基准数据集:COCO-4C和Food-4C,均包含四类不同的分组标准及对应的簇标签。实验表明,X-Cluster能在多个数据集上有效揭示有意义的划分。最后,我们将其应用于真实场景,包括揭示文生图生成模型中的隐藏偏见以及分析社交媒体上的图像传播现象。
原文摘要 · Abstract (English)
In this work, we introduce and study the novel task of Open-ended Semantic Multiple Clustering (OpenSMC). Given a large, unstructured image collection, the goal is to automatically discover several, diverse semantic clustering criteria (e.g., Activity or Location) from the images, and subsequently organize them according to the discovered criteria, without requiring any human input. Our framework, X-Cluster: eXploratory Clustering, treats text as a reasoning proxy: it concurrently scans the entire image collection, proposes candidate criteria in natural language, and groups images into meaningful clusters per criterion. This radically differs from previous works, which either assume predefined clustering criteria or fixed cluster counts. To evaluate X-Cluster, we create two new benchmarks, COCO-4C and Food-4C, each annotated with four distinct grouping criteria and corresponding cluster labels. Experiments show that X-Cluster can effectively reveal meaningful partitions on several datasets. Finally, we use X-Cluster to achieve various real-world applications, including uncovering hidden biases in text-to-image (T2I) generative models and analyzing image virality on social media. Project page: https://oatmealliu.github.io/xcluster.html
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。