实时驱动虚拟人说话,唇形同步更准,帧率超140,适合落地应用
Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching
- 基于流匹配生成,解决传统方法唇动不准和动作漂移问题
- 在HDTF数据集上达到8.50的唇同步置信度,每秒处理141帧
- 单张A10显卡即可实现0.17秒端到端延迟,适合实时场景
我们提出Livatar,一个实时音频驱动的虚拟人脸视频生成框架。现有方法普遍存在唇形不同步和长期姿态漂移的问题。本文采用基于流匹配的生成框架,结合系统优化,在HDTF数据集上实现了8.50的唇同步置信度,单张A10 GPU下达到141 FPS的吞吐量,端到端延迟仅为0.17秒。该系统使高保真虚拟人广泛应用于实时交互场景成为可能。项目代码与演示见https://www.hedra.com/及https://h-liu1997.github.io/Livatar-1/
原文摘要 · Abstract (English)
We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework. Coupled with system optimizations, Livatar achieves competitive lip-sync quality with a 8.50 LipSync Confidence on the HDTF dataset, and reaches a throughput of 141 FPS with an end-to-end latency of 0.17s on a single A10 GPU. This makes high-fidelity avatars accessible to broader applications. Our project is available at https://www.hedra.com/ with with examples at https://h-liu1997.github.io/Livatar-1/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。