从文本生成3D人物物体交互,先学通用动作再加对象特有细节。
EigenActor: Variant Body-Object Interaction Generation Evolved from Invariant Action Basis Reasoning

- 先提取通用动作模式,再添加对象特有交互风格。
- 在三个数据集上显著提升语义一致性和交互真实感。
- 适合做动作生成与跨模态交互研究的开发者参考。
本文研究从文本指令生成3D人-物交互(HOI)的跨模态合成任务。现有方法直接将文本映射到特定物体的3D动作,因跨模态差距大而性能受限。我们发现:相同交互意图(如“举起”)在不同物体(如椅子或杯子)上,其身体动作模式相似,仅交互风格不同。因此,学习有效的动作先验与对象交互先验对模型表现至关重要。为此,提出新策略:先推断无对象依赖的规范动作,再基于此补充对象特有交互风格。第一阶段学习类内共享的动作先验,将文本语义映射为动作特定的规范3D姿态;第二阶段学习对象功能,丰富具体交互样式。大量实验表明,该系统在三个大规模数据集上优于当前最优方法,显著提升语义一致性和交互真实性。
原文摘要 · Abstract (English)
This paper explores a cross-modality synthesis task that infers 3D human-object interactions (HOIs) from a given text-based instruction. Existing text-to-HOI synthesis methods mainly deploy a direct mapping from texts to object-specific 3D body motions, which may encounter a performance bottleneck since the huge cross-modality gap. In this paper, we observe that those HOI samples with the same interaction intention toward different targets, e.g., "lift a chair" and "lift a cup", always encapsulate similar action-specific body motion patterns while characterizing different object-specific interaction styles. Thus, learning effective action-specific motion priors and object-specific interaction priors is crucial for a text-to-HOI model and dominates its performances on text-HOI semantic consistency and body-object interaction realism. In light of this, we propose a novel body pose generation strategy for the text-to-HOI task: infer object-agnostic canonical body action first and then enrich object-specific interaction styles. Specifically, the first canonical body action inference stage focuses on learning intra-class shareable body motion priors and mapping given text-based semantics to action-specific canonical 3D body motions. Then, in the object-specific interaction inference stage, we focus on object affordance learning and enrich object-specific interaction styles on an inferred action-specific body motion basis. Extensive experiments verify that our proposed text-to-HOI synthesis system significantly outperforms other SOTA methods on three large-scale datasets with better semantic consistency and interaction realism performances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。