arXiv:2503.23121cs.CV2025-03中稿 · ICME 2025被引 2

用高效模型实现人体关节级交互建模,生成更真实的人物物体互动。

Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation

  • 双分支结构分别处理时空交互信息与文本物体特征融合
  • 仅需5%推理时间,显著优于以往方法且保持高精度
  • 适合需要实时生成复杂交互的视频/游戏应用

我们提出一种新型文本引导下人-物交互(HOI)生成方法,可在计算高效的前提下实现显式的关节级交互建模。此前方法将整个人体视为单一标记,难以捕捉精细的关节级交互,导致生成结果不自然;而将每个关节单独作为标记则会产生超过20倍的标记数量,大幅增加计算开销。为解决此问题,我们设计了高效显式关节级交互模型(EJIM),包含双分支HOI Mamba,分别高效建模时空交互信息;以及双分支条件注入器,将文本语义与物体几何融入人体与物体运动中。此外,我们引入动态交互模块与渐进掩码机制,迭代过滤无关关节,确保交互建模准确且细腻。在多个公开数据集上的大量定量与定性评估表明,EJIM在性能上远超现有方法,同时仅需5%的推理时间。代码已开源。

原文摘要 · Abstract (English)

We propose a novel approach for generating text-guided human-object interactions (HOIs) that achieves explicit joint-level interaction modeling in a computationally efficient manner. Previous methods represent the entire human body as a single token, making it difficult to capture fine-grained joint-level interactions and resulting in unrealistic HOIs. However, treating each individual joint as a token would yield over twenty times more tokens, increasing computational overhead. To address these challenges, we introduce an Efficient Explicit Joint-level Interaction Model (EJIM). EJIM features a Dual-branch HOI Mamba that separately and efficiently models spatiotemporal HOI information, as well as a Dual-branch Condition Injector for integrating text semantics and object geometry into human and object motions. Furthermore, we design a Dynamic Interaction Block and a progressive masking mechanism to iteratively filter out irrelevant joints, ensuring accurate and nuanced interaction modeling. Extensive quantitative and qualitative evaluations on public datasets demonstrate that EJIM surpasses previous works by a large margin while using only 5\% of the inference time. Code is available \href{https://github.com/Huanggh531/EJIM}{here}.

人物交互高效建模文本生成视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。