arXiv:2410.15866cs.CV2024-10ECCV

用CLIP特征识别电影中的20类视觉母题,准确率达91%。

Visual Motif Identification: Elaboration of a Curated Comparative Dataset and Classification Methods

  • 基于CLIP特征+浅层网络,设计分类模型
  • 测试集上F1-score达0.91,效果出色
  • 适合影视分析与艺术风格研究者

在电影中,视觉母题是具有艺术或美学意义的反复出现的图像构图。其在视觉艺术与媒体历史中的运用对研究者和电影人颇具吸引力。本文旨在通过提出一种新机器学习模型,实现对这些母题的识别与分类。我们利用自建数据集,展示如何通过浅层网络结合特定损失函数,基于CLIP模型提取的特征,将图像分为20类不同母题,取得令人惊喜的效果:在测试集上达到0.91的F1-score。此外,我们还进行了多项消融实验,验证所选输入特征、模型结构与超参数的合理性。

原文摘要 · Abstract (English)

In cinema, visual motifs are recurrent iconographic compositions that carry artistic or aesthetic significance. Their use throughout the history of visual arts and media is interesting to researchers and filmmakers alike. Our goal in this work is to recognise and classify these motifs by proposing a new machine learning model that uses a custom dataset to that end. We show how features extracted from a CLIP model can be leveraged by using a shallow network and an appropriate loss to classify images into 20 different motifs, with surprisingly good results: an $F_1$-score of 0.91 on our test set. We also present several ablation studies justifying the input features, architecture and hyperparameters used.

视觉母题CLIP图像分类电影分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。