用3D高斯泼溅实现多轮对话中真实且有社交意识的虚拟人脸生成
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation
- 结合3D网格驱动与3D高斯泼溅,实现高效高保真人脸渲染
- 在自建数据集上达到最先进的真实感与社交感知性能
- 适合虚拟社交、数字人对话等需考虑人际关系的场景
说话头生成在虚拟现实尤其是多轮对话社交场景中日益重要。现有方法存在明显局限:基于网格的3D方法可建模双人对话但纹理不真实;大模型驱动的2D方法虽外观自然,但计算成本过高。最近基于3D高斯泼溅(3DGS)的方法实现了高效且逼真的渲染,但仍仅支持单人且忽略社交关系。本文提出RSATalker,首个利用3DGS实现多轮对话中真实且具备社交感知能力的说话头生成框架。方法首先从语音驱动3D网格面部运动,再将3D高斯分布绑定至网格面片,生成高保真2D虚拟人视频。为捕捉人际动态,设计社交感知模块,通过可学习查询机制将血缘/非血缘、平等/不平等关系编码为高层嵌入。构建三阶段训练范式,并建立包含语音-网格-图像三元组及社交关系标注的RSATalker数据集。大量实验表明,RSATalker在真实感与社交感知方面均达到当前最优水平。代码与数据集将公开。
原文摘要 · Abstract (English)
Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack realistic textures, while large-model-based 2D methods produce natural appearances but incur prohibitive computational costs. Recently, 3D Gaussian Splatting (3DGS) based methods achieve efficient and realistic rendering but remain speaker-only and ignore social relationships. We introduce RSATalker, the first framework that leverages 3DGS for realistic and socially-aware talking head generation with support for multi-turn conversation. Our method first drives mesh-based 3D facial motion from speech, then binds 3D Gaussians to mesh facets to render high-fidelity 2D avatar videos. To capture interpersonal dynamics, we propose a socially-aware module that encodes social relationships, including blood and non-blood as well as equal and unequal, into high-level embeddings through a learnable query mechanism. We design a three-stage training paradigm and construct the RSATalker dataset with speech-mesh-image triplets annotated with social relationships. Extensive experiments demonstrate that RSATalker achieves state-of-the-art performance in both realism and social awareness. The code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。