通过语义对齐与时间重排序,提升视频配乐推荐准确率
Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking

- 先全局语义匹配,再细粒度时间对齐候选音乐
- R@10从14.2升至18.3,中位排名从75降至46
- 适合需要精准配乐的视频创作与内容推荐场景
我们提出VTMR,一种两阶段视频配乐推荐框架。第一阶段在联合音视频文本表征空间中对齐视频与音乐信号,利用粗粒度全局嵌入高效检索语义匹配的候选音乐;第二阶段通过关注视频与音乐的时间序列,重排候选结果,捕捉细粒度时间对应关系。在视频配乐推荐任务上,多模态检索阶段将R@10从14.2提升至15.9,中位排名从75降至58;时间重排序器进一步将R@10提升至18.3,中位排名降至46,体现更丰富查询编码与时间对齐的互补增益。人工偏好测试表明,VTMR在整体偏好上与商业基线相当,且优于生成基线的音乐质量。
原文摘要 · Abstract (English)
We present VTMR, a two-stage framework for Video-To-Music Recommendation. In Stage~1, VTMR aligns comprehensive video and music signals in a joint audio-visual-text representation space and efficiently retrieves semantically compatible candidates using coarse global embeddings. In Stage~2, it reranks the retrieved candidates by attending to the temporal sequences of both video and music, thereby capturing fine-grained temporal correspondence. Evaluated on the video-to-music recommendation task, the multimodal retrieval stage improves R@10 from 14.2 to 15.9 and Median Rank from 75 to 58 over the strongest baseline; the temporal reranker further boosts R@10 to 18.3 and Median Rank to 46, demonstrating complementary gains from richer query encoding and temporal alignment. A human preference study confirms that VTMR is on par with a commercial baseline in overall preference, while outperforming a generative baseline in music quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。