用生成式AI在流媒体中智能修复冗余视频,提升画质不增带宽。
End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
- 服务端编码+客户端生成修复,端到端优化视频流质量。
- 最高提升11 VMAF分,画质显著改善且无需额外带宽。
- 模块化设计适配新算法,适合研究者和流媒体开发者。
视频流传输的核心挑战在于高质量与流畅播放之间的平衡。传统编解码器虽已优化此权衡,但因无法利用上下文信息,必须传输完整视频数据。本文提出ELVIS(端到端学习型视频流增强管道),结合服务端编码优化与客户端生成修复技术,实现冗余视频数据的删除与重建。其模块化架构可集成不同编解码器、修复模型及质量评估指标,具备良好可扩展性。实验表明,当前技术可在基准上实现最高11 VMAF分的提升,但实时应用仍受计算开销制约。ELVIS为将生成式AI引入视频流管道奠定了基础,实现画质提升而无需增加带宽。
原文摘要 · Abstract (English)
The primary challenge of video streaming is to balance high video quality with smooth playback. Traditional codecs are well tuned for this trade-off, yet their inability to use context means they must encode the entire video data and transmit it to the client. This paper introduces ELVIS (End-to-end Learning-based VIdeo Streaming Enhancement Pipeline), an end-to-end architecture that combines server-side encoding optimizations with client-side generative in-painting to remove and reconstruct redundant video data. Its modular design allows ELVIS to integrate different codecs, inpainting models, and quality metrics, making it adaptable to future innovations. Our results show that current technologies achieve improvements of up to 11 VMAF points over baseline benchmarks, though challenges remain for real-time applications due to computational demands. ELVIS represents a foundational step toward incorporating generative AI into video streaming pipelines, enabling higher quality experiences without increased bandwidth requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。