用视觉基础模型实现高保真、低带宽视频流,抗网络波动。
Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model
- 基于视觉基础模型,端到端生成视频流
- 比H.265节省62.5%带宽,保持相近画质
- 支持实时传输,适合弱网环境
视频流媒体是互联网的基础服务,但在带宽受限或偏远地区,画质难以保障。传统像素编码方案已接近极限,而新兴的神经增强或生成式流媒体在延迟和画质上仍不足。受视觉基础模型(VFM)成功启发,本文提出Morphe,首个基于VFM的端到端生成式视频流范式。通过在模拟网络约束下联合训练视觉分词器与可变分辨率时空优化,结合智能丢包机制构建鲁棒系统。实验表明,Morphe在挑战性网络环境下实现实时、抗丢包视频传输,相比H.265节省62.5%带宽,同时保持相近视觉质量,标志着VFM赋能多媒体流媒体的重要里程碑。
原文摘要 · Abstract (English)
Video streaming is a fundamental Internet service, while the quality still cannot be guaranteed especially in poor network conditions such as bandwidth-constrained and remote areas. Existing works mainly work towards two directions: traditional pixel-codec streaming nearly approaches its limit and is hard to step further in compression; the emerging neural-enhanced or generative streaming usually fall short in latency and visual fidelity, hindering their practical deployment. Inspired by the recent success of vision foundation model (VFM), we strive to harness the powerful video understanding and processing capacities of VFM to achieve generalization, high fidelity and loss resilience for real-time video streaming with even higher compression rate. We present the first revolutionized paradigm that enables VFM-based end-to-end generative video streaming towards this goal. Specifically, Morphe employs joint training of visual tokenizers and variable-resolution spatiotemporal optimization under simulated network constraints. Additionally, a robust streaming system is constructed that leverages intelligent packet dropping to resist real-world network perturbations. Extensive evaluation demonstrates that Morphe achieves comparable visual quality while saving 62.5\% bandwidth compared to H.265, and accomplishes real-time, loss-resilient video delivery in challenging network environments, representing a milestone in VFM-enabled multimedia streaming solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。