在5G边缘云上实现低延迟音视频语音增强,提升交互体验。
Audio-Visual Speech Enhancement: Architectural Design and Deployment Strategies
- 音视频融合架构:用CNN+OpenCV+LSTM实现同步增强。
- 5G和有线网络满足实时性要求,压缩可降80%数据量。
- 适合开发低延迟多媒体服务的工程师参考。
实时音视频语音增强(AVSE)是沉浸式交互媒体服务的关键,但受网络延迟、上行带宽和计算延迟制约。本文设计并部署了一个基于公共5G边缘网络的完整云-边协同AVSE系统,结合基于CNN的声学增强、OpenCV的面部特征提取与LSTM融合网络以保持时间一致性,运行于Vodafone兼容的AWS Wavelength边缘云。通过大量压力测试,分析了不同网络负载与自适应多媒体配置下的端到端性能。结果表明,计算部署在边缘对满足实时一致性至关重要,且上行带宽常为交互式AVSE服务的主要瓶颈。仅5G和有线以太网能持续满足未压缩音视频块的通信延迟要求;而激进压缩可降低80%数据量且感知质量损失微小,使系统在受限条件下仍稳定运行。进一步揭示处理延迟与增强质量间存在根本权衡:模型简化可降低延迟,但在低信噪比场景下会损害重建性能。研究显示,当网络与计算资源协调得当时,公共5G边缘环境可支撑实时交互式AVSE任务,尽管性能余量小于专用基础设施。本研究的架构洞见为新兴5G边缘云平台上的延迟敏感型多媒体与感知增强服务设计提供了实用指导。
原文摘要 · Abstract (English)
Real-time audio-visual speech enhancement (AVSE) is a key enabler for immersive and interactive multimedia services, yet its performance is tightly constrained by network latency, uplink capacity, and computational delay. This paper presents the design, deployment, and evaluation of a complete cloud-edge-assisted AVSE system operating over a public 5G edge network. The system integrates CNN-based acoustic enhancement and OpenCV-based facial feature extraction with an LSTM fusion network to preserve temporal coherence, and is deployed on a Vodafone-compatible AWS Wavelength edge cloud. Through extensive stress testing, we analyze end-to-end performance under varying network load and adaptive multimedia profiles. Results show that compute placement at the network edge is critical for meeting real-time coherence constraints, and that uplink capacity is often the dominant bottleneck for interactive AVSE services. Only 5G and wired Ethernet consistently satisfied the required communication delay bound for uncompressed audio-video chunks, while aggressive compression reduced payload sizes by up to 80% with negligible perceptual degradation, enabling robust operation under constrained conditions. We further demonstrate a fundamental trade-off between processing latency and enhancement quality, where reduced model complexity lowers delay but degrades reconstruction performance in low-SNR scenarios. Our findings indicate that public 5G edge environments can sustain real-time, interactive AVSE workloads when network and compute resources are carefully orchestrated, although performance margins remain tighter than in dedicated infrastructures. The architectural insights derived from this study provide practical guidelines for the design of delay-sensitive multimedia and perceptual enhancement services on emerging 5G edge-cloud platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。