用深度学习压缩4D视频,让复杂三维内容轻松进入混合现实设备。
Computer Vision and Deep Learning for 4D Augmented Reality
- 用深度学习构建4D视频的紧凑表示,保留形状与外观细节。
- 成功将高精度3D表演捕获导入HoloLens等混合现实设备。
- 适合关注XR内容压缩与实时渲染的研究者和开发者。
扩展现实(XR)平台中的4D视频前景广阔,为人类与计算机交互及媒体感知带来全新方式。本文证明了在微软混合现实平台中渲染4D视频的可行性,使我们能够相对轻松地将来自CVSSP的任意3D表演捕获移植至如HoloLens这样的XR设备。然而,当3D模型过于复杂、包含数百万个顶点时,当前硬件与通信系统面临严重的数据带宽瓶颈。为此,本项目开发了一种基于深度学习的紧凑表示方法,有效学习4D视频序列的形状与外观特征,并实现无损重建,显著降低传输开销。
原文摘要 · Abstract (English)
The prospect of 4D video in Extended Reality (XR) platform is huge and exciting, it opens a whole new way of human computer interaction and the way we perceive the reality and consume multimedia. In this thesis, we have shown that feasibility of rendering 4D video in Microsoft mixed reality platform. This enables us to port any 3D performance capture from CVSSP into XR product like the HoloLens device with relative ease. However, if the 3D model is too complex and is made up of millions of vertices, the data bandwidth required to port the model is a severe limitation with the current hardware and communication system. Therefore, in this project we have also developed a compact representation of both shape and appearance of the 4d video sequence using deep learning models to effectively learn the compact representation of 4D video sequence and reconstruct it without affecting the shape and appearance of the video sequence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。