用Whisper提升语音特征提取速度,让虚拟主播更实时逼真。
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
- 用Whisper替换传统语音特征提取模型,减少延迟
- 在三数据集上验证,处理更快且生成质量更高
- 适合需要实时互动的AI面试培训场景
本文研究实时虚拟人像生成在面试培训中的应用,聚焦音频特征提取(AFE)带来的延迟问题。为解决此挑战,提出并实现一个全集成系统,将传统AFE模型替换为OpenAI的Whisper,利用其编码器优化处理流程,提升整体效率。在三个公开数据集上对两个开源实时模型进行评估,结果表明Whisper不仅加速了处理过程,还提升了渲染质量的特定方面,使生成的虚拟人像更具真实感和响应性。该系统显著增强了沉浸式、交互式培训的效果,拓展了基于AI的虚拟形象在面试训练中的应用潜力。
原文摘要 · Abstract (English)
This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsiveness in real-time applications. To address these issues, we propose and implement a fully integrated system that replaces conventional AFE models with Open AI's Whisper, leveraging its encoder to optimize processing and improve overall system efficiency. Our evaluation of two open-source real-time models across three different datasets shows that Whisper not only accelerates processing but also improves specific aspects of rendering quality, resulting in more realistic and responsive talking-head interactions. These advancements make the system a more effective tool for immersive, interactive training applications, expanding the potential of AI-driven avatars in interviewer training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。