TinyMusician让手机端生成高质量音乐成为可能
TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- 用知识蒸馏和混合精度量化压缩大模型
- 模型大小减半,性能保留93%以上
- 适合移动端音乐生成,无需依赖云端
生成模型在音乐生成领域取得突破性进展,基于Transformer的架构已成性能新标杆。然而其大规模参数带来巨大算力需求和推理延迟,难以部署于智能手机、可穿戴设备等边缘设备。本文提出TinyMusician,从MusicGen(当前最优音乐生成模型)中蒸馏出轻量级模型。创新点包括:(i) 阶段性混合双向与偏斜KL散度,(ii) 自适应混合精度量化。实验表明,TinyMusician仅需55%的模型尺寸,即可保持MusicGen-Small 93%的性能。它是首个无需云端支持、可在移动端部署且保持高音质的音乐生成模型。
原文摘要 · Abstract (English)
The success of the generative model has gained unprecedented attention in the music generation area. Transformer-based architectures have set new benchmarks for model performance. However, their practical adoption is hindered by some critical challenges: the demand for massive computational resources and inference time, due to their large number of parameters. These obstacles make them infeasible to deploy on edge devices, such as smartphones and wearables, with limited computational resources. In this work, we present TinyMusician, a lightweight music generation model distilled from MusicGen (a State-of-the-art music generation model). TinyMusician integrates two innovations: (i) Stage-mixed Bidirectional and Skewed KL-Divergence and (ii) Adaptive Mixed-Precision Quantization. The experimental results demonstrate that TinyMusician retains 93% of the MusicGen-Small performance with 55% less model size. TinyMusician is the first mobile-deployable music generation model that eliminates cloud dependency while maintaining high audio fidelity and efficient resource usage
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。