arXiv:2503.19901cs.CV2025-03CVPR被引 70

用任务标记统一控制人类与场景互动,实现多技能灵活切换。

TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization

  • 将任务和本体感知分别建模为标记,通过掩码融合统一策略。
  • 支持可变输入长度,使技能可适配新场景和复杂任务。
  • 适合需要多技能集成的动画生成与具身智能研究者。

生成多样且物理合理的真人-场景交互(HSI)对计算机动画和具身AI至关重要。尽管已有进展,当前方法多为针对特定任务设计独立控制器,难以整合多种技能应对复杂任务(如携带物体坐下)。为此,我们提出TokenHSI,一种基于Transformer的统一策略,实现多技能融合与灵活适应。核心思路是将人形本体感知建模为共享标记,通过掩码机制与不同任务标记结合。该统一策略促进技能间知识共享,支持多任务训练。此外,模型支持可变长度输入,可灵活适配新场景。通过训练额外任务标记器,不仅能修改交互目标几何形状,还可协调多个技能完成复杂任务。实验表明,该方法在多种HSI任务中显著提升泛化性、适应性和可扩展性。

原文摘要 · Abstract (English)

Synthesizing diverse and physically plausible Human-Scene Interactions (HSI) is pivotal for both computer animation and embodied AI. Despite encouraging progress, current methods mainly focus on developing separate controllers, each specialized for a specific interaction task. This significantly hinders the ability to tackle a wide variety of challenging HSI tasks that require the integration of multiple skills, e.g., sitting down while carrying an object. To address this issue, we present TokenHSI, a single, unified transformer-based policy capable of multi-skill unification and flexible adaptation. The key insight is to model the humanoid proprioception as a separate shared token and combine it with distinct task tokens via a masking mechanism. Such a unified policy enables effective knowledge sharing across skills, thereby facilitating the multi-task training. Moreover, our policy architecture supports variable length inputs, enabling flexible adaptation of learned skills to new scenarios. By training additional task tokenizers, we can not only modify the geometries of interaction targets but also coordinate multiple skills to address complex tasks. The experiments demonstrate that our approach can significantly improve versatility, adaptability, and extensibility in various HSI tasks. Website: https://liangpan99.github.io/TokenHSI/

人机交互生成模型多技能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。