arXiv:2601.13148cs.CVcs.HC2026-01IJCV被引 3

打造可实时对话的逼真3D虚拟人,支持语音驱动面部动画。

ICo3D: An Interactive Conversational 3D Virtual Human

  • 用高斯点云构建可动的3D人脸与人体模型
  • 通过语音驱动实现口型与表情精准同步
  • 适合元宇宙、虚拟助手等沉浸式交互场景

本文提出Interactive Conversational 3D Virtual Human(ICo3D),一种生成可交互、可对话且逼真的3D虚拟人方法。基于多视角捕捉的主体数据,构建可动画化的3D人脸模型与动态3D身体模型,均采用高斯点云(splatting Gaussian primitives)渲染。两者融合后形成适用于实时交互的拟真人形象。为实现对话能力,引入大语言模型(LLM)。对话过程中,虚拟人语音作为驱动信号,精确控制人脸模型动画,确保音视频同步。提出改进方案:SWinGS++用于提升身体重建质量,HeadGaS++优化人脸重建效果,并提供无伪影的头身模型融合方法。系统展示多个实时对话应用场景。该方法实现了语音与文字双模交互,适用于游戏、虚拟助教、个性化教育等沉浸式应用领域。

原文摘要 · Abstract (English)

This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi-view captures of a subject, we create an animatable 3D face model and a dynamic 3D body model, both rendered by splatting Gaussian primitives. Once merged together, they represent a lifelike virtual human avatar suitable for real-time user interactions. We equip our avatar with an LLM for conversational ability. During conversation, the audio speech of the avatar is used as a driving signal to animate the face model, enabling precise synchronization. We describe improvements to our dynamic Gaussian models that enhance photorealism: SWinGS++ for body reconstruction and HeadGaS++ for face reconstruction, and provide as well a solution to merge the separate face and body models without artifacts. We also present a demo of the complete system, showcasing several use cases of real-time conversation with the 3D avatar. Our approach offers a fully integrated virtual avatar experience, supporting both oral and written form interactions in immersive environments. ICo3D is applicable to a wide range of fields, including gaming, virtual assistance, and personalized education, among others. Project page: https://ico3d.github.io/

虚拟人3D生成对话系统高斯渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。