arXiv:2606.10550cs.CVcs.GR2026-06

用单目视频生成可实时立体显示的虚拟头像,无需特殊设备。

LentiAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic Communication

论文配图:LentiAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic Communication
图 1 · 摘自论文原文
  • 用自然转头模拟多视角监督,提升单目重建细节。
  • 实现32视角实时渲染,4K全息屏下保持10.65帧率。
  • 适合远程会议、沉浸式通信场景,无需佩戴眼镜。

实时立体视频通信长期受限于专用拍摄设备或仅能呈现单视角肖像。本文提出LentiAvatar,一种基于高斯头像的系统,将单目肖像捕捉与无镜片子像素编码透镜屏结合,实现实时自动立体通信。从单目视频中,系统重建可控头像,并针对显示屏产生的横向视区进行优化。利用自然头部转动作为伪多视角(PMV)监督,约束单目训练中弱观测区域,包括头发、耳朵、下颌轮廓和颈部边界。可靠侧视图通过偏航分箱、对齐虚拟相机,并在严格头与发域内监督;轮廓感知损失与分阶段正则化进一步抑制鬼影、透明度泄漏和深度不稳定,同时保留横向细节。运行时,系统渲染32个虚拟视角,并以校准的子像素路由掩码编码为4K透镜光栅。实时追踪原型维持10.65 FPS,特定用户压缩驱动器使同一显示管线达到38.49 FPS。

原文摘要 · Abstract (English)

Real-time stereoscopic video communication has long been a goal of immersive telepresence, yet practical systems still require specialized capture rigs or reduce remote users to a single portrait view. We present LentiAvatar, a Gaussian head-avatar system that connects monocular avatar capture with subpixel-encoded glasses-free lenticular display for real-time autostereoscopic communication. From a monocular portrait video, LentiAvatar reconstructs a controllable head avatar and optimizes it for the lateral viewing zones induced by the display. The method uses natural head turns as pseudo-multiview (PMV) supervision to constrain regions that are otherwise weakly observed in monocular training, including hair, ears, jaw contours, and neck boundaries. Reliable side frames are yaw-binned, aligned to virtual cameras, and supervised within a strict head-and-hair domain; contour-aware losses and staged regularization further suppress ghosting, alpha leakage, and depth instability while preserving lateral detail. At runtime, LentiAvatar renders 32 virtual views and encodes them into a 4K lenticular raster with calibrated subpixel-routing masks. The live-tracker prototype sustains 10.65 FPS, and a subject-specific distilled driver raises the same display pipeline to 38.49 FPS.

立体显示虚拟头像实时渲染单目重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。