用海报图文融合提升电影类型分类准确率
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
- 通过OCR提取海报文字,结合视觉与文本特征进行跨注意力融合
- 在13882张IMDb海报上测试,性能优于多个主流模型
- 适合电影推荐、营销分析等场景使用
电影海报不仅是装饰,更是精心设计以体现影片类型、情节和氛围的关键载体。长期以来,海报广泛用于影院、广告牌及数字屏幕,为观众提供影片的预览信息。电影类型分类在影视营销、用户参与和推荐系统中具有重要作用。以往研究多聚焦于剧情摘要、字幕、预告片和影片片段,而海报能在上映前激发公众兴趣,提供关键线索。本文提出一种基于视觉与文本双模态协同的跨注意力融合框架,解决多标签电影类型分类问题。首先通过OCR提取海报中的文字并获取其嵌入表示;随后引入跨注意力融合模块,动态分配视觉与文本嵌入的权重。实验基于来自互联网电影数据库(IMDb)的13882张海报验证框架效果,结果表明该模型表现优异,超越多个主流架构。
原文摘要 · Abstract (English)
Movie posters are not just decorative; they are meticulously designed to capture the essence of a movie, such as its genre, storyline, and tone/vibe. For decades, movie posters have graced cinema walls, billboards, and now our digital screens as a form of digital posters. Movie genre classification plays a pivotal role in film marketing, audience engagement, and recommendation systems. Previous explorations into movie genre classification have been mostly examined in plot summaries, subtitles, trailers and movie scenes. Movie posters provide a pre-release tantalizing glimpse into a film's key aspects, which can ignite public interest. In this paper, we presented the framework that exploits movie posters from a visual and textual perspective to address the multilabel movie genre classification problem. Firstly, we extracted text from movie posters using an OCR and retrieved the relevant embedding. Next, we introduce a cross-attention-based fusion module to allocate attention weights to visual and textual embedding. In validating our framework, we utilized 13882 posters sourced from the Internet Movie Database (IMDb). The outcomes of the experiments indicate that our model exhibited promising performance and outperformed even some prominent contemporary architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。