构建首个视频音乐对齐数据集,助力多模态理解
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
- 人工与机器协同标注4.8万组视频音乐对,覆盖节奏、情感等多维度
- 提出多维评估框架,量化视频与音乐在节奏、情绪、主题上的对齐度
- 适合研究跨模态对齐、内容生成的学者与开发者使用
本文提出HarmonySet,一个用于推动视频-音乐理解的综合性数据集。该数据集包含48,328组多样化的视频-音乐配对,每对均标注了节奏同步性、情感一致性、主题连贯性及文化相关性等详细信息。我们设计了一种多步骤人机协同标注框架,结合人类洞察与机器生成描述,识别关键转折点并评估多维度对齐情况。此外,提出了新的评估框架,包含任务与度量标准,用于衡量视频与音乐在节奏、情感、主题和文化背景等方面的多维对齐。大量实验表明,HarmonySet及其配套评估框架显著提升了多模态模型捕捉和分析视频与音乐复杂关系的能力。
原文摘要 · Abstract (English)
This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance. We propose a multi-step human-machine collaborative framework for efficient annotation, combining human insights with machine-generated descriptions to identify key transitions and assess alignment across multiple dimensions. Additionally, we introduce a novel evaluation framework with tasks and metrics to assess the multi-dimensional alignment of video and music, including rhythm, emotion, theme, and cultural context. Our extensive experiments demonstrate that HarmonySet, along with the proposed evaluation framework, significantly improves the ability of multimodal models to capture and analyze the intricate relationships between video and music.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。