arXiv:2608.27123cs.CV2026-08

实时直播中人物视频编辑,保持表情一致且低延迟。

EditaLive! Unified Character Video Editing for Live Streaming

论文配图:EditaLive! Unified Character Video Editing for Live Streaming
图 1 · 摘自论文原文
  • 基于预训练动画模型,通过参考帧编辑实现指令驱动的人像视频生成。
  • 采用因果流式生成与对齐自回放蒸馏,压缩为两步采样器,支持实时推理。
  • 适用于需要高保真表情和低延迟的直播互动场景,如虚拟主播、在线会议。

传统视频编辑侧重场景级内容,而直播更关注人物主体。然而,直接应用现有编辑方法于以人物为中心的直播仍具挑战,因其易导致面部表情不一致,且通常依赖多个离线推理步骤,难以实现实时交互。本文提出 EditLive,一种面向实时直播的人物视频编辑新框架。我们从预训练图像动画模型 Wan-Animate 出发,其天然解耦外观与运动,通过参考帧编辑与视频重建方式,利用收集的 CharEdit-50K 数据集将其改造为基于指令的人像视频编辑基础模型。此外,我们将模型从离线双向生成转为因果流式生成,并设计对齐自回放蒸馏策略,将模型压缩为两步采样器:固定 RoPE 与对齐强制减少训练-推理差异,首帧保留的稀疏注意力过滤冗余历史信息,缓解外观漂移。大量实验表明,EditLive 在保持面部表情真实性的前提下,实现了领先水平的编辑性能与低延迟实时推理。

原文摘要 · Abstract (English)

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.

视频编辑直播实时生成表情一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。