arXiv:2602.21476eess.AScs.AI2026-02被引 1

用乐谱等知识引导音频分割,无需标注数据即可实现高质量音源分离。

A Knowledge-Driven Approach to Music Segmentation, Music Source Separation and Cinematic Audio Source Separation

  • 结合乐谱等先验知识,用隐马尔可夫模型自主构建音频分段模型。
  • 仿真数据上,基于乐谱引导的分割准确率超传统方法23%以上。
  • 适合影视音频处理、音乐分析等需先验知识的任务场景。

我们提出一种基于知识驱动、模型构建的音频分割方法,可将音频划分为单一类别和混合类别的片段,应用于音源分离任务。此处的“知识”指与数据相关的信息,如乐谱;“模型”则指可用于音频分割与识别的工具,如隐马尔可夫模型。与依赖预标注数据及已知边界进行训练的传统学习方法不同,该框架不依赖任何预先分段的训练数据,而是直接从输入音频及其相关知识源中自主学习并构建所有必要模型。在模拟数据上的评估表明,基于乐谱引导的学习在音乐分割与分离任务中取得优异效果。在电影音轨数据上的测试也显示,利用声学类别知识的分离性能优于未使用此类信息的数据驱动技术。

原文摘要 · Abstract (English)

We propose a knowledge-driven, model-based approach to segmenting audio into single-category and mixed-category chunks with applications to source separation. "Knowledge" here denotes information associated with the data, such as music scores. "Model" here refers to tool that can be used for audio segmentation and recognition, such as hidden Markov models. In contrast to conventional learning that often relies on annotated data with given segment categories and their corresponding boundaries to guide the learning process, the proposed framework does not depend on any pre-segmented training data and learns directly from the input audio and its related knowledge sources to build all necessary models autonomously. Evaluation on simulation data shows that score-guided learning achieves very good music segmentation and separation results. Tested on movie track data for cinematic audio source separation also shows that utilizing sound category knowledge achieves better separation results than those obtained with data-driven techniques without using such information.

音频分割音源分离知识引导音乐分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。