Hallo2实现4K高清、时长数十分钟的音频驱动人脸动画,支持文本控制。
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
- 通过图像空间增强与噪声补丁丢弃,解决长时间生成中的外观漂移问题。
- 首次实现4K分辨率下持续数十分钟的音频驱动人脸视频生成。
- 支持文本标签调节表情,提升生成内容的可控性与多样性。
基于潜在扩散模型的人脸图像动画近期取得显著进展,如Hallo在短时视频合成上表现优异。本文提出Hallo2,对原方法进行多项改进:首先,扩展至生成长时视频,针对外观漂移和时间伪影问题,在条件运动帧的图像空间引入含高斯噪声的补丁丢弃策略,增强视觉一致性和时间连贯性;其次,实现4K分辨率人脸视频生成,通过潜码向量量化与时间对齐技术保持时序一致性,并结合高质量解码器达成4K输出;第三,引入可调语义文本标签作为条件输入,超越传统音频驱动,提升生成可控性与内容多样性。据我们所知,Hallo2是首个实现4K分辨率、长达数小时且由音频与文本共同驱动的人脸动画方法。我们在公开数据集HDTF、CelebV及自建“Wild”数据集上进行了大量实验,结果表明该方法在长时人脸视频动画任务中达到当前最优性能,成功生成高保真、可调控的4K级内容,持续时长可达数十分钟。
原文摘要 · Abstract (English)
Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities. First, we extend the method to produce long-duration videos. To address substantial challenges such as appearance drift and temporal artifacts, we investigate augmentation strategies within the image space of conditional motion frames. Specifically, we introduce a patch-drop technique augmented with Gaussian noise to enhance visual consistency and temporal coherence over long duration. Second, we achieve 4K resolution portrait video generation. To accomplish this, we implement vector quantization of latent codes and apply temporal alignment techniques to maintain coherence across the temporal dimension. By integrating a high-quality decoder, we realize visual synthesis at 4K resolution. Third, we incorporate adjustable semantic textual labels for portrait expressions as conditional inputs. This extends beyond traditional audio cues to improve controllability and increase the diversity of the generated content. To the best of our knowledge, Hallo2, proposed in this paper, is the first method to achieve 4K resolution and generate hour-long, audio-driven portrait image animations enhanced with textual prompts. We have conducted extensive experiments to evaluate our method on publicly available datasets, including HDTF, CelebV, and our introduced "Wild" dataset. The experimental results demonstrate that our approach achieves state-of-the-art performance in long-duration portrait video animation, successfully generating rich and controllable content at 4K resolution for duration extending up to tens of minutes. Project page https://fudan-generative-vision.github.io/hallo2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。