压缩语音说话人分离模型,平衡效率与性能,适用于医疗调度等实时场景。
Efficiency-Performance Trade-offs in Neural Speaker Diarization via Structured Pruning and Low-Bit Quantization
- 通过结构化剪枝和低比特量化压缩模型,降低内存占用。
- FP16下模型大小减半,实时因子几乎不变,相对DER增加40%。
- 发现低延迟会显著降效,缓冲非总是有益,适合资源受限设备部署。
流式说话人分离对时间敏感的医疗调度至关重要,但在资源受限硬件上部署需更小更快的模型。我们使用SIMSAMU数据集(模拟医疗调度对话)评估压缩前的流式行为。在不同流式延迟预算下分析性能,发现额外缓冲并非始终有益,极低延迟运行点会显著降低性能。研究显示模型压缩以牺牲性能为代价换取更小内存占用;其中,使用FP16可使模型大小减半,实时因子基本不变,但相对错误率(DER)相比基线上升40%。本工作刻画了实时部署中的效率-性能权衡,推动语音技术在时间关键场景中实现可靠人机通信。
原文摘要 · Abstract (English)
Streaming speaker diarization is crucial for time-critical medical dispatch, but deploying it on resource-constrained hardware requires smaller, faster models. Using SIMSAMU, a dataset of simulated medical-dispatch conversations, we evaluate streaming behavior before compressing the segmentation model with pruning and low-bit quantization. We characterize performance across a range of streaming latency budgets and find that additional buffering is not consistently beneficial, while very low-latency operating points can substantially degrade performance. Our study shows that model compression trades performance for memory footprint, and we highlight an operating point where FP16 reduces model size by half with essentially unchanged real-time factor, at a cost of a 40\% relative DER increase against the baseline. This work characterizes the trade-offs for real-time deployment and contributes to speech technology that can enable reliable human communication in time-critical contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。