arXiv:2606.09595cs.IR2026-06

构建可配置的多模态电影推荐基准,验证不同视觉证据的有效性。

Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation

论文配图:Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation
图 1 · 摘自论文原文
  • 整合影片、预告片与缩略图的多模态特征,统一配置管理。
  • 缩略图模型提供强而可扩展的物品表征,效果优于预告片。
  • 适合研究多模态推荐融合策略与视觉信号差异的研究者。

电影是长时序音视频内容,但现有推荐基准常依赖预告片、缩略图或元数据。这些数据在语义和可扩展性上存在差异:完整影片保留消费级证据,预告片聚焦宣传亮点,缩略图则提供稀疏但全量的视觉信号。本文提出 Popcorn,一个可配置的多模态电影推荐视觉证据基准,融合与片名对齐的完整影片/预告片嵌入,以及通过现代视觉与视觉-语言模型编码的 MovieLens 关联缩略图特征。Popcorn 通过单一配置契约标准化了模态组合、融合、划分、评估及 LLM 增强元数据流程。实验表明,缩略图的 VLM 模型提供了强大且可扩展的物品侧证据;而控制变量的预告片与完整影片对比显示,不同视觉证据源不可互换:选择来源与融合策略显著影响排序准确率、覆盖率、多样性与校准性。代码已开源:https://github.com/RecSys-lab/Popcorn。

原文摘要 · Abstract (English)

Movies are long-form audiovisual works, yet recommender benchmarks often rely on trailers, thumbnails, or metadata. These sources differ in semantics and scalability: full movies preserve consumption-level evidence, trailers concentrate promotional highlights, and thumbnails provide sparse but catalog-scale visual signals. We present Popcorn, a configurable benchmark for visual evidence in multimodal movie recommendation, combining title-aligned full-movie/trailer embeddings with MovieLens-linked thumbnail features encoded by modern visual and vision-language models. Popcorn standardizes modality assembly, fusion, splitting, evaluation, and LLM-augmented metadata through a single configuration contract. Experiments show that thumbnail VLMs provide strong, scalable item-side evidence, while controlled trailer/full-movie comparisons show that visual evidence sources are not interchangeable: the choice of source and fusion strategy affects ranking accuracy, coverage, diversity, and calibration. The framework is available at https://github.com/RecSys-lab/Popcorn.

多模态推荐视觉证据可配置基准电影推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。