arXiv:2410.23730cs.CV2024-10SIGGRAPH被引 8

GPT-4V零样本评估时尚审美,表现接近人类但对同色系穿搭排序不佳

An Empirical Analysis of GPT-4V's Performance on Fashion Aesthetic Evaluation

  • 首次零样本测试GPT-4V在时尚审美评估中的表现
  • 预测结果与人类判断基本一致,相关性较高
  • 在相似颜色搭配的排序上表现较差,存在视觉感知局限

时尚审美评估是估算图像中人物所穿服装是否适合其个人形象的任务。本文首次系统考察了GPT-4V在该任务上的零样本表现。实验结果表明,其预测结果在我们的数据集上与人类判断具有较好的一致性;同时发现,当面对颜色相近的穿搭时,模型在排序方面存在明显困难。代码已公开于https://github.com/st-tech/gpt4v-fashion-aesthetic-evaluation。

原文摘要 · Abstract (English)

Fashion aesthetic evaluation is the task of estimating how well the outfits worn by individuals in images suit them. In this work, we examine the zero-shot performance of GPT-4V on this task for the first time. We show that its predictions align fairly well with human judgments on our datasets, and also find that it struggles with ranking outfits in similar colors. The code is available at https://github.com/st-tech/gpt4v-fashion-aesthetic-evaluation.

视觉理解时尚评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。