用统一语义占据表示3D场景,生成更自然的人体动作。
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
- 用双向三平面分解压缩场景语义占据表示
- 在ShapeNet、PROX和Replica数据集上表现领先
- 适合需要精细场景理解的动作生成任务
人体动作在3D场景中的合成严重依赖场景理解,而现有方法主要关注场景结构,忽视语义信息。本文提出一种基于统一场景语义占据(SSO)的运动合成框架——SSOMotion。设计双向三平面分解以获得紧凑的SSO表示,并通过CLIP编码与共享线性降维将场景语义映射至统一特征空间,既保留细粒度语义结构,又大幅减少冗余计算。进一步结合指令推导的运动方向与帧级场景查询实现动作控制。在包含家具的杂乱场景(ShapeNet)、PROX及Replica扫描数据集上的大量实验与消融研究验证了该方法的卓越性能、有效性和泛化能力。代码将公开于https://github.com/jingyugong/SSOMotion。
原文摘要 · Abstract (English)
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability. Code will be publicly available at https://github.com/jingyugong/SSOMotion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。