无需模板,用可逆网络高效重建动态场景中物体的语义3D形态。
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene
- 用可逆神经网络预测线性混合皮肤权重,简化动态物体变形建模。
- 支持任意刚体、非刚体交互,单视角或稀疏视图下实现高精度重建。
- 相比传统LBS方法训练更快,适合复杂动态场景的实时语义重建。
尽管神经隐式模型在3D表面重建方面取得进展,但处理任意刚体、非刚体或可变形实体之间相互作用的动态环境仍具挑战。通用重建方法通常需深度图、光流或预训练图像特征才能获得合理结果,且多依赖潜在编码捕捉逐帧形变。另一类方法为特定实体(如人体)设计,依赖模板模型。相比之下,部分无模板方法虽避免上述限制,采用传统线性混合皮肤(LBS)权重以实现精细变形表示,但优化过程复杂,训练时间长。为此,本文提出TFS-NeRF:一种无需模板的动态场景语义3D NeRF,适用于从稀疏或单视角RGB视频中重建含双及以上实体交互的场景,且训练效率优于其他基于LBS的方法。框架利用可逆神经网络(INN)预测LBS权重,简化训练流程。通过解耦交互实体的运动并分别优化各实体的皮肤权重,有效生成准确且语义可分的几何结构。大量实验表明,该方法在复杂交互下能高质量重建可变形与不可变形物体,同时相比现有方法具备更优训练效率。
原文摘要 · Abstract (English)
Despite advancements in Neural Implicit models for 3D surface reconstruction, handling dynamic environments with interactions between arbitrary rigid, non-rigid, or deformable entities remains challenging. The generic reconstruction methods adaptable to such dynamic scenes often require additional inputs like depth or optical flow or rely on pre-trained image features for reasonable outcomes. These methods typically use latent codes to capture frame-by-frame deformations. Another set of dynamic scene reconstruction methods, are entity-specific, mostly focusing on humans, and relies on template models. In contrast, some template-free methods bypass these requirements and adopt traditional LBS (Linear Blend Skinning) weights for a detailed representation of deformable object motions, although they involve complex optimizations leading to lengthy training times. To this end, as a remedy, this paper introduces TFS-NeRF, a template-free 3D semantic NeRF for dynamic scenes captured from sparse or single-view RGB videos, featuring interactions among two entities and more time-efficient than other LBS-based approaches. Our framework uses an Invertible Neural Network (INN) for LBS prediction, simplifying the training process. By disentangling the motions of interacting entities and optimizing per-entity skinning weights, our method efficiently generates accurate, semantically separable geometries. Extensive experiments demonstrate that our approach produces high-quality reconstructions of both deformable and non-deformable objects in complex interactions, with improved training efficiency compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。