arXiv:2601.22127cs.CVcs.GR2026-01被引 2

用音频控制视频,精准编辑真人说话头像的口型和内容。

EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers

  • 基于扩散变换器,通过音频条件实现视频到视频的精准编辑。
  • 支持添加、删除和重定时口型内容,保持动作连贯性与唇形同步。
  • 适合影视后期、虚拟主播等需要精细控制说话视频的场景。

当前生成式视频模型擅长根据文本或图像生成新内容,但在编辑已有录制视频方面存在明显短板:微调台词时需保持动作、时间连贯性、说话人身份一致及准确的唇形同步。我们提出 EditYourself,一种基于扩散变换器(DiT)的音频驱动视频到视频(V2V)编辑框架,支持基于字幕对说话头像视频进行修改,包括无缝添加、移除和重定时视觉上的语音内容。该框架在通用视频扩散模型基础上,引入音频条件与区域感知、编辑导向的训练扩展,实现精确的唇形同步与时空一致的重构,通过时空修复合成真实的人体动作,同时在长时间内维持视觉保真度与身份一致性。本工作为生成式视频模型成为专业视频后期制作工具迈出关键一步。

原文摘要 · Abstract (English)

Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserving motion, temporal coherence, speaker identity, and accurate lip synchronization. We introduce EditYourself, a DiT-based framework for audio-driven video-to-video (V2V) editing that enables transcript-based modification of talking head videos, including the seamless addition, removal, and retiming of visually spoken content. Building on a general-purpose video diffusion model, EditYourself augments its V2V capabilities with audio conditioning and region-aware, edit-focused training extensions. This enables precise lip synchronization and temporally coherent restructuring of existing performances via spatiotemporal inpainting, including the synthesis of realistic human motion in newly added segments, while maintaining visual fidelity and identity consistency over long durations. This work represents a foundational step toward generative video models as practical tools for professional video post-production.

视频编辑扩散模型语音同步说话头像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。