用深度学习分离短视频背景音乐,还原原始音轨解决版权问题
Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline
- 结合音频分离与跨模态检索,从混音中提取原始音轨
- 在2万段音频上实现高精度去背景音乐,160个视频验证恢复效果
- 适合平台方处理侵权内容,也可用于内容版权保护
YouTube Shorts、TikTok等短视频平台面临严峻的版权合规挑战,侵权者常嵌入任意背景音乐(BGM)以掩盖原声轨道(OST),逃避原创性检测。为此,我们提出一种融合音乐源分离(MSS)与跨模态视频-音乐检索(CMVMR)的新型流水线,可有效分离任意BGM并恢复原始OST,保障视频音频真实性。为支持该研究,我们构建了两个领域专用数据集:OASD-20K包含20,000段混合BGM与OST音频片段;OSVAR-160是首个专为短视频修复设计的基准数据集,含1,121对视频与混音音频样本。实验表明,该流水线不仅能高精度去除任意背景音乐,还可成功还原原始音轨,确保内容完整性。该方法为用户生成内容中的版权问题提供了一种伦理且可扩展的解决方案。
原文摘要 · Abstract (English)
Short video platforms like YouTube Shorts and TikTok face significant copyright compliance challenges, as infringers frequently embed arbitrary background music (BGM) to obscure original soundtracks (OST) and evade content originality detection. To tackle this issue, we propose a novel pipeline that integrates Music Source Separation (MSS) and cross-modal video-music retrieval (CMVMR). Our approach effectively separates arbitrary BGM from the original OST, enabling the restoration of authentic video audio tracks. To support this work, we introduce two domain-specific datasets: OASD-20K for audio separation and OSVAR-160 for pipeline evaluation. OASD-20K contains 20,000 audio clips featuring mixed BGM and OST pairs, while OSVAR-160 is a unique benchmark dataset comprising 1,121 video and mixed-audio pairs, specifically designed for short video restoration tasks. Experimental results demonstrate that our pipeline not only removes arbitrary BGM with high accuracy but also restores OSTs, ensuring content integrity. This approach provides an ethical and scalable solution to copyright challenges in user-generated content on short video platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。