用多模态预训练特征精准分类电影预告片类型
Movie Trailer Genre Classification Using Multimodal Pretrained Features
- 融合视觉、音频、文本等多模态预训练特征
- 在MovieNet数据集上达到更高精度与mAP
- 适合对视频理解与跨模态学习感兴趣的研究者
我们提出一种新颖的电影类型分类方法,利用多种现成的预训练模型提取视觉场景、物体、人物、文本、语音、音乐和音效等高层特征。为智能融合这些特征,我们训练了轻量级分类器,具有低时延和低内存需求。采用Transformer模型,直接使用预告片全部视频与音频帧,无需时间池化,充分捕捉各元素间的对应关系,突破传统方法仅用固定少量帧的限制。该方法融合来自不同任务与模态、不同维度、不同时间长度及复杂依赖关系的特征,优于现有先进模型在精确率、召回率和平均精度均值(mAP)上的表现。为促进后续研究,我们公开了整个MovieNet数据集的预训练特征、分类代码及训练好的模型。
原文摘要 · Abstract (English)
We introduce a novel method for movie genre classification, capitalizing on a diverse set of readily accessible pretrained models. These models extract high-level features related to visual scenery, objects, characters, text, speech, music, and audio effects. To intelligently fuse these pretrained features, we train small classifier models with low time and memory requirements. Employing the transformer model, our approach utilizes all video and audio frames of movie trailers without performing any temporal pooling, efficiently exploiting the correspondence between all elements, as opposed to the fixed and low number of frames typically used by traditional methods. Our approach fuses features originating from different tasks and modalities, with different dimensionalities, different temporal lengths, and complex dependencies as opposed to current approaches. Our method outperforms state-of-the-art movie genre classification models in terms of precision, recall, and mean average precision (mAP). To foster future research, we make the pretrained features for the entire MovieNet dataset, along with our genre classification code and the trained models, publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。