实时生成人像视频特效,支持交互式编辑且保持身份一致
StreamingEffect: Real-Time Human-Centric Video Effect Generation
- 采用上下文视频编辑架构,通过双向教师模型蒸馏出轻量因果自回归学生模型
- 仅需4步采样即可实现720p高清视频实时编辑,单H200 GPU部署
- 构建13万条人像特效数据集,支持在线关键帧控制与流式交互
实时人像视频特效生成对电商直播、娱乐和Vlog等应用极具价值,但受限于数据稀缺与可部署编辑模型不足。该任务需在保持人物身份、背景内容与时间一致性前提下,实现实时视频到视频的表达性特效添加。现有加速研究多集中于文本到视频生成,而高效视频编辑蒸馏仍鲜有探索。本文提出StreamingEffect框架,采用上下文视频编辑架构,训练高质量双向教师模型,并将其蒸馏为因果自回归学生模型,将采样步骤从50步降至4步。引入关键帧控制机制,支持在线注入参考特效帧并沿流传播以实现交互式编辑。为解决数据瓶颈,构建了迄今为止最大的人像视频特效数据集VideoEffect-130K,包含70K特效视频与60K编辑视频,覆盖600类特效,数据源自短视频与剪辑平台。实验表明,该方法可在单张H200 GPU上实现720p分辨率的实时高质量视频编辑。
原文摘要 · Abstract (English)
Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet remains difficult due to the lack of suitable data and deployable editing models. Unlike generic video generation, this task requires real-time video-to-video editing that adds expressive effects while preserving human identity, background content, and temporal consistency. Existing acceleration efforts mainly focus on text-to-video generation, while efficient distillation for video editing remains largely underexplored. In this paper, we present \textbf{StreamingEffect}, a real-time human-centric streaming video effect framework. We adopt an in-context video editing architecture and train a high-quality bidirectional teacher, then distill it into a causal autoregressive student and further reduce sampling from 50 steps to 4 steps. We also introduce keyframe control, allowing reference effect frames to be injected online and propagated through the stream for interactive editing. To address the data bottleneck, we construct \textbf{VideoEffect-130K}, to our knowledge the largest human-centric video effect dataset, containing 70K effect videos and 60K editing videos across 600 effect categories curated from short-video and editing platforms. Experiments show that our method enables real-time, high-quality 720p video editing on a single H200 GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。