用事件相机+视频帧,让快速运动的人体3D重建更清晰
ExFMan: Rendering 3D Dynamic Humans with Hybrid Monocular Blurry Frames and Events
- 结合事件数据与视频帧,按运动速度自适应调整损失权重
- 提出速度感知光度损失和速度相关事件损失,提升模糊区域精度
- 适合做虚拟角色、动作捕捉的开发者或研究人员使用
近年来,随着神经渲染技术的发展,单目视频中动态人体的三维重建取得了显著进展,广泛应用于虚拟现实中的虚拟角色创建。然而,在真实场景中,快速人体运动(如跑步、跳舞)常导致视频出现运动模糊,使得重建结果在形状和外观上存在明显不一致,尤其在手部、腿部等高速运动部位。本文提出ExFMan,首个利用帧图像与生物启发式事件相机混合数据进行快速运动人体高质量渲染的神经渲染框架。核心思想是互补利用事件数据的高时间分辨率信息,并根据重建人体各区域的运动速度自适应调节RGB帧与事件数据的损失权重,有效缓解了运动模糊带来的不一致性。具体地,先在标准空间构建人体速度场并映射至图像空间以定位模糊区域;随后设计两种新损失:速度感知光度损失与速度相关事件损失,引导双模态优化。此外,引入新的姿态正则化与透明度损失,以实现连续姿态与清晰边界。在合成与真实数据集上的大量实验表明,ExFMan能显著提升人体重建的清晰度与质量。
原文摘要 · Abstract (English)
Recent years have witnessed tremendous progress in the 3D reconstruction of dynamic humans from a monocular video with the advent of neural rendering techniques. This task has a wide range of applications, including the creation of virtual characters for virtual reality (VR) environments. However, it is still challenging to reconstruct clear humans when the monocular video is affected by motion blur, particularly caused by rapid human motion (e.g., running, dancing), as often occurs in the wild. This leads to distinct inconsistency of shape and appearance for the rendered 3D humans, especially in the blurry regions with rapid motion, e.g., hands and legs. In this paper, we propose ExFMan, the first neural rendering framework that unveils the possibility of rendering high-quality humans in rapid motion with a hybrid frame-based RGB and bio-inspired event camera. The ``out-of-the-box'' insight is to leverage the high temporal information of event data in a complementary manner and adaptively reweight the effect of losses for both RGB frames and events in the local regions, according to the velocity of the rendered human. This significantly mitigates the inconsistency associated with motion blur in the RGB frames. Specifically, we first formulate a velocity field of the 3D body in the canonical space and render it to image space to identify the body parts with motion blur. We then propose two novel losses, i.e., velocity-aware photometric loss and velocity-relative event loss, to optimize the neural human for both modalities under the guidance of the estimated velocity. In addition, we incorporate novel pose regularization and alpha losses to facilitate continuous pose and clear boundary. Extensive experiments on synthetic and real-world datasets demonstrate that ExFMan can reconstruct sharper and higher quality humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。