用深度学习实现零延迟自动混音,专为现场演出设计
AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
- 基于深度学习构建端到端系统,实时处理现场录音串扰
- 实现零延迟混音,在16个通道上保持音频视觉同步
- 可扩展预测增益外的其他混音参数,适配未来升级
本文提出一种面向现场演出的深度学习自动多轨混音系统。现场演出中,声道常受邻近乐器声泄漏干扰,且音视频同步至关重要,对音频延迟有严格限制。本工作主要解决输入声道中声泄漏问题,实现零延迟混音。尽管近年已有自动混音研究进展,但多数集中在离线制作和独立音源信号处理,据我们所知,这是首个专为现场演出设计的端到端深度学习混音系统。当前系统仅预测单声道增益,但其架构结合过往研究成果,可轻松扩展至未来预测其他关键混音参数。
原文摘要 · Abstract (English)
In this work, we present a deep learning-based automatic multitrack music mixing system catered towards live performances. In a live performance, channels are often corrupted with acoustic bleeds of co-located instruments. Moreover, audio-visual synchronization is of critical importance thus putting a tight constraint on the audio latency. In this work we primarily tackle these two challenges of handling bleeds in the input channels to produce the music mix with zero latency. Although there have been several developments in the field of automatic music mixing in recent times, most or all previous works focus on offline production for isolated instrument signals and to the best of our knowledge, this is the first end-to-end deep learning system developed for live music performances. Our proposed system currently predicts mono gains for a multitrack input, but its design along with the precedent set in past works, allows for easy adaptation to future work of predicting other relevant music mixing parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。