针对带宽受限下的多源语音传输,提出感知与任务协同的分布式编码方法。
Task and Perception-aware Distributed Source Coding for Correlated Speech under Bandwidth-constrained Channels
- 基于神经主成分分析的分布式编码,动态适应比特率
- 任务感知损失使重建语音在低带宽下提升52%性能
- 适用于实时语音传输的AR/VR场景,兼顾真实感与任务精度
新兴无线AR/VR应用需在不可靠、带宽受限的信道上,从多个资源受限设备实时传输相关高保真语音。现有基于自编码器的语音源编码方法无法同时满足:(1) 不重训练下的动态码率自适应,(2) 利用多语音源间的相关性,(3) 平衡下游任务损失与重建语音的真实性。本文提出一种基于神经主成分分析(NDPCA)的分布式源编码算法,用于多源语音向中心接收端传输。该方法包含感知感知任务损失函数,平衡感知真实感与任务性能。实验表明,在无任务感知(19%)和任务感知(52%)设置下,相比朴素自编码器方法,均实现显著的PSNR提升。在低带宽场景下,性能逼近理论上限(所有源送入单编码器)。此外,我们提供了速率-失真-感知权衡曲线,支持根据应用需求自适应选择。
原文摘要 · Abstract (English)
Emerging wireless AR/VR applications require real-time transmission of correlated high-fidelity speech from multiple resource-constrained devices over unreliable, bandwidth-limited channels. Existing autoencoder-based speech source coding methods fail to address the combination of the following - (1) dynamic bitrate adaptation without retraining the model, (2) leveraging correlations among multiple speech sources, and (3) balancing downstream task loss with realism of reconstructed speech. We propose a neural distributed principal component analysis (NDPCA)-aided distributed source coding algorithm for correlated speech sources transmitting to a central receiver. Our method includes a perception-aware downstream task loss function that balances perceptual realism with task-specific performance. Experiments show significant PSNR improvements under bandwidth constraints over naive autoencoder methods in task-agnostic (19%) and task-aware settings (52%). It also approaches the theoretical upper bound, where all correlated sources are sent to a single encoder, especially in low-bandwidth scenarios. Additionally, we present a rate-distortion-perception trade-off curve, enabling adaptive decisions based on application-specific realism needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。