分离建模人体与场景,通过映射机制实现高精度动态交互渲染
Dynamic Avatar-Scene Rendering from Human-centric Context
- 分步建模人体与场景,再通过统一变换函数融合
- 在单目视频上实现更清晰的人体-场景边界渲染
- 适合需要精细人景交互的虚拟角色生成应用
从单目视频中重建动态人物与真实环境的交互是一个重要且具挑战性的任务。尽管4D神经渲染取得了显著进展,现有方法或整体建模动态场景,或分开建模场景与背景并引入参数化人体先验。前者忽略场景各组件(尤其是人体)的独特运动特性,导致重建不完整;后者忽视分离建模组件间的信息交互,造成空间不一致和人景边界视觉伪影。为此,我们提出“分离-映射”(StM)策略,引入专用信息映射机制连接独立定义与优化的模型。方法采用共享变换函数对每个高斯属性进行统一处理,避免了冗余的成对交互,提升计算效率的同时确保人体与其周围环境的空间与视觉一致性。在单目视频数据集上的大量实验表明,StM在视觉质量和渲染精度上显著优于现有最先进方法,尤其在复杂的人景交互边界表现突出。
原文摘要 · Abstract (English)
Reconstructing dynamic humans interacting with real-world environments from monocular videos is an important and challenging task. Despite considerable progress in 4D neural rendering, existing approaches either model dynamic scenes holistically or model scenes and backgrounds separately aim to introduce parametric human priors. However, these approaches either neglect distinct motion characteristics of various components in scene especially human, leading to incomplete reconstructions, or ignore the information exchange between the separately modeled components, resulting in spatial inconsistencies and visual artifacts at human-scene boundaries. To address this, we propose {\bf Separate-then-Map} (StM) strategy that introduces a dedicated information mapping mechanism to bridge separately defined and optimized models. Our method employs a shared transformation function for each Gaussian attribute to unify separately modeled components, enhancing computational efficiency by avoiding exhaustive pairwise interactions while ensuring spatial and visual coherence between humans and their surroundings. Extensive experiments on monocular video datasets demonstrate that StM significantly outperforms existing state-of-the-art methods in both visual quality and rendering accuracy, particularly at challenging human-scene interaction boundaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。