arXiv:2607.28032cs.CV2026-07

单图生成实时可驱动的高保真头像,无需外部追踪

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

论文配图:Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
图 1 · 摘自论文原文
  • 分两轴解耦:计算轴内化驱动,特征轴拆分面部区域建模
  • 单卡运行全流程推理速度最快,性能超越现有方法
  • 适合实时数字人、虚拟主播场景使用

从单张图像生成逼真可驱动的头部虚拟人仍是数字人合成中的核心挑战。尽管近期3D高斯溅射方法已取得良好效果,但其依赖外部追踪系统,且推理延迟未计入测量;同时采用统一表示,使几何上不同的面部区域相互纠缠,限制了表现力与渲染质量。我们提出SpiD(Split and Drive),一种基于双解耦轴的单图高斯头部虚拟人框架。计算轴将每帧驱动内化,消除推理时对外部追踪的依赖;特征轴将虚拟人分解为三个专用高斯分支,分别建模几何上独立的面部区域。大量实验表明,相比现有最优方法,该方法在性能上持续领先,且在包含完整驱动流程的情况下,单卡推理速度最快。

原文摘要 · Abstract (English)

Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.

3D生成高斯溅射数字人实时驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。