端到端学习声波生成,让水下视频传输更抗干扰且实时可用。
E2E-WAVE: End-to-End Learned Waveform Generation for Underwater Video Multicasting

- 将语义相似性直接嵌入声波信号,错误时优先保留相近内容。
- 在2.3 kbps下实现16帧/秒128×128视频,比最强基线提升5 dB PSNR。
- 单通道训练即可泛化到未知水下环境,适合低带宽水下通信场景。
我们提出E2E-WAVE,首个面向水下视频多播的端到端学习波形生成系统。声学信道比特误码率高达20%–46%,前向纠错(FEC)反而加剧错误——当超过解码阈值后,LDPC会增加而非减少误码。E2E-WAVE通过将语义相似性直接嵌入物理层波形,在不可避免解码错误时优先选择语义相近的词元,而非任意错误替换。结合VideoGPT分词(压缩率达1024倍)与可训练波形库及全可微分OFDM传输,E2E-WAVE在较优水下信道(NOF1)中实现比最强FEC保护基线高出+5 dB(19.26%)PSNR和+0.10(14.28%)SSIM;同时在2.3 kbps信道上实现实时16 FPS、128×128分辨率视频传输——传统数字调制无法实现。在更恶劣信道(BCH1、NCS1)中性能差距进一步扩大。仅在单一信道训练即可泛化至未见水下环境,无需重新训练;而HEVC在低于5 kbps时失效,SoftCast的高斯白噪声假设在频率选择性信道中崩溃。
原文摘要 · Abstract (English)
We present E2E-WAVE, the first end-to-end learned waveform generation system for underwater video multicasting. Acoustic channels exhibit 20--46% bit error rates where forward error correction becomes counterproductive -- LDPC increases rather than decreases errors beyond its decoding threshold. E2E-WAVE addresses this by embedding semantic similarity directly into physical layer waveforms: when decoding errors are unavoidable, the system preferentially selects semantically similar tokens rather than arbitrary corruption. Combining VideoGPT tokenization (1024x compression) with a trainable waveform bank and fully differentiable OFDM transmission, E2E-WAVE achieves +5 dB (19.26%) PSNR and +0.10 (14.28%) SSIM over the strongest FEC-protected baseline in less challenging underwater channel (NOF1) while delivering real-time 16 FPS video at 128x128 resolution over 2.3 kbps channels -- impossible for conventional digital modulation. The performance gap only increases in harsher channels (BCH1, NCS1). Trained on a single channel, E2E-WAVE generalizes to unseen underwater environments without retraining, while HEVC fails at sub-5 kbps rates and SoftCast's AWGN assumptions collapse on frequency-selective channels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。