基于预训练模型的视频情绪分类网页工具,支持自定义视频分析。
VEMOCLAP: A video emotion classification web application
- 融合视频帧与音频的预训练特征,用多头交叉注意力高效融合。
- 在Ekman-6数据集上准确率提升4.3%,达当前最优水平。
- 开源可在线使用,支持上传本地视频或输入YouTube链接。
我们提出VEMOCLAP:一种基于预训练特征的视频情绪分类网络应用,是首个可直接访问且开源的网页工具,能分析任意用户提供的视频情感内容。该方法改进了先前工作,利用开源预训练模型处理视频帧与音频,并通过多头交叉注意力机制高效融合特征。实验表明,在Ekman-6视频情绪数据集上,分类准确率提升4.3%,达到当前最佳性能。系统提供在线服务,用户可上传本地视频或输入YouTube链接运行模型。欢迎访问 serkansulun.com/app 体验。
原文摘要 · Abstract (English)
We introduce VEMOCLAP: Video EMOtion Classifier using Pretrained features, the first readily available and open-source web application that analyzes the emotional content of any user-provided video. We improve our previous work, which exploits open-source pretrained models that work on video frames and audio, and then efficiently fuse the resulting pretrained features using multi-head cross-attention. Our approach increases the state-of-the-art classification accuracy on the Ekman-6 video emotion dataset by 4.3% and offers an online application for users to run our model on their own videos or YouTube videos. We invite the readers to try our application at serkansulun.com/app.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。