边云协同的语音情感描述生成,提升效率与隐私保护。
Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
- 轻量边缘模型本地初稿,高不确定词块上传云端验证。
- 在MER2024上BLEU最高提升62.7%,延迟降低1.4倍。
- 适合注重隐私与实时性的边缘语音应用开发者。
语音情感描述(SEC)利用大型音视频语言模型从语音中生成丰富、上下文感知的情感描述。然而,由于资源受限的边缘设备面临巨大的计算压力,且传输生物特征音频存在隐私风险,实际部署仍具挑战。尽管小型音视频语言模型可实现设备端高效SEC,但其容量有限,常削弱细微的副语言建模与细粒度情感定位能力。本文提出基于不确定性引导推测解码(UGSD)的边云协同框架:轻量级边缘模型本地生成草稿,仅将高不确定性词块选择性上传至更强的云端验证器进行校验。在MER2024基准上的实验表明,该方法显著提升BLEU值,最高达62.7%。相比纯边缘模型,UGSD还实现1.4倍更低延迟与8.5倍更高令牌吞吐量。这些结果实证刻画了可部署SEC系统中质量-效率-隐私之间的权衡关系。
原文摘要 · Abstract (English)
Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains challenging due to the substantial computational demands on resource-constrained edge devices and the privacy risks of transmitting biometric audio. While smaller audio-language models enable efficient on-device SEC, their limited capacity often weakens subtle paralinguistic modeling and fine-grained affective grounding. We propose an edge-cloud collaborative framework based on Uncertainty-Guided Speculative Decoding (UGSD). A lightweight edge model drafts captions locally, and only high-uncertainty token blocks are selectively escalated to a stronger cloud verifier for validation. Experiments on the MER2024 benchmark demonstrate substantial BLEU improvements up to 62.7%. UGSD further achieves 1.4x lower latency and 8.5x higher token throughput compared to an edge-only model. These results empirically characterize the quality-efficiency-privacy trade-off in deployable SEC systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。