无需训练即可实现音频驱动的通用人脸视频编辑,支持表情与口型同步调整。
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
- 通过统一动画控制机制,用初始参考帧实现无需训练的动态编辑
- 口型同步准确率高,支持不同头部姿态和表情的灵活控制
- 适合需要快速生成个性化视频内容的创作者或应用开发人员
人脸视频编辑旨在根据音频或视频流修改人脸视频的特定属性。以往方法通常局限于唇部区域重演,或需训练专用模型提取关键点以实现动作迁移。本文提出一种无需训练的通用人脸视频编辑框架,支持基于首参考帧变化的外观编辑、基于语音变化的口型编辑,或两者结合。该框架基于统一动画控制(UAC)机制,利用源图像逆向隐空间编码,实现视觉驱动的形变控制、音频驱动的说话控制及帧间时序控制。通过调整初始参考帧,可适应不同场景,实现对特定头部旋转和面部表情的精细编辑。实验表明,该方法在口型编辑任务中实现了更精确且同步的唇动,在外观编辑任务中具备更强的运动迁移灵活性。演示地址:https://alice01010101.github.io/RASA/
原文摘要 · Abstract (English)
Portrait video editing focuses on modifying specific attributes of portrait videos, guided by audio or video streams. Previous methods typically either concentrate on lip-region reenactment or require training specialized models to extract keypoints for motion transfer to a new identity. In this paper, we introduce a training-free universal portrait video editing framework that provides a versatile and adaptable editing strategy. This framework supports portrait appearance editing conditioned on the changed first reference frame, as well as lip editing conditioned on varied speech, or a combination of both. It is based on a Unified Animation Control (UAC) mechanism with source inversion latents to edit the entire portrait, including visual-driven shape control, audio-driven speaking control, and inter-frame temporal control. Furthermore, our method can be adapted to different scenarios by adjusting the initial reference frame, enabling detailed editing of portrait videos with specific head rotations and facial expressions. This comprehensive approach ensures a holistic and flexible solution for portrait video editing. The experimental results show that our model can achieve more accurate and synchronized lip movements for the lip editing task, as well as more flexible motion transfer for the appearance editing task. Demo is available at https://alice01010101.github.io/RASA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。