解耦语义与几何,生成更真实的真人-场景互动动作。
DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction

- 分离动作意图与空间约束,实现精准控制。
- 减少47%的穿透错误,提升29%语义一致性。
- 适合虚拟人、动画生成和交互系统开发者。
生成符合物理规律的人-场景交互(HSI)仍是计算机视觉与虚拟角色开发中的关键挑战。尽管现有生成模型能合成多样动作,但存在语义-几何纠缠的归纳偏见:训练数据中空间约束常与特定动作强相关,导致模型在面对严格几何线索时忽略语义意图,引发身体穿模等物理幻觉。为此,本文提出DeSeG,一种分层框架,显式解耦语义意图与几何约束。首先,引入残差语义规划器,将文本指令与标准化目标体素编码至紧凑隐空间,实现与轨迹无关的细粒度语义控制。其次,提出物理正则化扩散执行器,将可微排斥势场直接嵌入扩散目标,强制生成防碰撞动作。在Lingo数据集上的实验表明,DeSeG达到当前最优性能,相比基线平均场景穿模减少47%,语义对齐提升29%。
原文摘要 · Abstract (English)
Synthesizing physically plausible human-scene interactions (HSI) remains a critical challenge in computer vision and the development of human avatars. Although recent generative models enable diverse motion synthesis, they suffer from an inductive bias referred to as semantic-geometric entanglement. Because spatial constraints often strongly correlate with specific actions in training data, monolithic models will learn the shortcut bias, aggressively overriding the semantic intent when faced with strict geometric cues. Furthermore, this entanglement exacerbates physical hallucinations, such as body-scene penetrations. To address these limitations, we propose DeSeG, a hierarchical framework that explicitly decouples semantic intent from geometric constraints. First, we introduce a Residual Semantic Planner that encodes textual instructions and canonicalized goal voxels into a compact latent space, enabling fine-grained semantic control independent of spatial trajectories. Second, we propose a physics regularized diffusion executor that incorporates differentiable repulsive potential fields directly into the diffusion objective, enforcing collision-aware motion generation. Extensive experiments on the Lingo dataset demonstrate that DeSeG achieves state-of-the-art performance, reducing mean scene penetration by 47% and improving semantic alignment by 29% over the SOTA baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。