arXiv:2607.26106eess.IVcs.CV2026-07

让视频提示流在丢包时仍能清晰还原,关键靠可分级提示排序。

ScalablePromptus: Scalable and High-Fidelity Prompt-Based Video Streaming

论文配图:ScalablePromptus: Scalable and High-Fidelity Prompt-Based Video Streaming
图 1 · 摘自论文原文
  • 用语义与色彩感知反推提示,提升重建质量
  • 通过球面插值生成中间帧,保持视频连贯性
  • 训练时模拟丢包,使提示支持任意截断重建

基于提示的视频流将紧凑的语义提示而非像素内容传输,实现超低码率通信。然而,现有 Promptus 框架对网络波动敏感,部分接收提示会导致质量灾难性下降。我们提出 ScalablePromptus,引入语义与色彩感知的提示反演、球面线性插值生成中间帧,并最关键的采用丢包训练策略,生成有序提示表示。这使得接收端可从任意截断提示中重建有意义视频,无需额外适配。在稳定网络下,性能略有提升;在丢包条件下,相比基线,性能下降减少82%-95%,使提示流具备实际部署可行性。

原文摘要 · Abstract (English)

Prompt-based video streaming transmits compact semantic prompts instead of pixel-level content for generative reconstruction, enabling ultra-low-bitrate communication. However, the state-of-the-art Promptus framework is vulnerable to network fluctuation, where partially received prompts lead to catastrophic quality collapse. We propose ScalablePromptus, which enhances Promptus with semantic and color-aware prompt inversion, spherical linear interpolation for intermediate frames, and--most critically--a dropout training strategy that produces rank-ordered prompt representations. This allows the receiver to reconstruct meaningful video from arbitrarily truncated prompts without any adaptation. Under stable networks, ScalablePromptus achieves modest quality gains. Under lossy conditions, it reduces the performance degradation caused by truncation by 82%-95% compared to the baseline, making prompt-based streaming robust enough for real-world deployment.

视频生成提示工程低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。