arXiv:2503.00371cs.CV2025-03被引 5

让3D人体动作生成与分析互相促进,提升真实感与语义一致性。

Jointly Understand Your Command and Intention:Reciprocal Co-Evolution between Scene-Aware 3D Human Motion Synthesis and Analysis

  • 构建合成-分析协同进化框架,双向优化动作生成与理解。
  • 通过分阶段生成实现目标推断、路径规划与姿态合成,增强动作合理性。
  • 适合关注3D动作生成、人-场景交互建模的研究者。

场景感知的人体动作生成与分析是两个紧密相关的任务,需要对3D身体动作、3D场景和文本描述等多种模态进行联合理解。本文提出协同进化合成-分析(CESA)流水线,将两者整合并相互促进。具体而言,基于文本的场景感知动作生成可从同一文本描述中生成多样化的室内动作样本,增强人-场景交互的类内多样性,显著提升鲁棒性动作分析系统的训练效果;反之,动作分析对每个生成样本施加语义审视,确保其与给定文本描述一致,从而提升动作生成的真实感。考虑到真实世界中的室内动作具有目标导向和路径引导特性,我们设计了分阶段生成策略,将文本驱动的场景特定动作生成分解为三个阶段:目标推断、路径规划与姿态合成。将CESA与该强大分阶段生成模型结合,共同提升了3D场景中真实人体动作的生成质量与鲁棒性分析能力。

原文摘要 · Abstract (English)

As two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, we integrate these two paired processes into a Co-Evolving Synthesis-Analysis (CESA) pipeline and mutually benefit their learning. Specifically, scene-aware text-to-human synthesis generates diverse indoor motion samples from the same textual description to enrich human-scene interaction intra-class diversity, thus significantly benefiting training a robust human motion analysis system. Reciprocally, human motion analysis would enforce semantic scrutiny on each synthesized motion sample to ensure its semantic consistency with the given textual description, thus improving realistic motion synthesis. Considering that real-world indoor human motions are goal-oriented and path-guided, we propose a cascaded generation strategy that factorizes text-driven scene-specific human motion generation into three stages: goal inferring, path planning, and pose synthesizing. Coupling CESA with this powerful cascaded motion synthesis model, we jointly improve realistic human motion synthesis and robust human motion analysis in 3D scenes.

3D动作生成人-场景交互协同学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。