用跨模态注意力提升场景感知的人体动作生成能力
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
- 通过双向跨模态注意力融合场景与人体特征
- 在复杂2D场景中生成合理的新姿态,准确率显著提升
- 适合做交互式机器人或虚拟人动作生成的研究者
人体可操作性学习旨在根据场景上下文预测合理的新型姿态,使估计的姿态在场景中具有实际可行性。该任务对机器感知和自动化交互导航至关重要,但因姿态与动作组合呈指数级增长,问题极具挑战。现有二维场景中的人体可操作性预测数据集与方法仍十分有限。本文提出一种新颖的互交叉注意力机制,通过双模态空间特征图的相互关注来编码场景上下文,用于可操作性预测。方法将任务解耦,降低复杂度:首先使用基于全局场景编码的变分自编码器(VAE)采样人体可能位置;接着利用局部上下文编码,通过分类器从已有姿态候选集中预测潜在姿态模板;随后,分别使用两个VAE,基于局部上下文和模板类别,采样姿态模板的尺度与形变参数。实验表明,该方法在复杂2D场景中的人体可操作性注入任务上显著优于先前基线。
原文摘要 · Abstract (English)
Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action within the scene. While the task is fundamental to machine perception and automated interactive navigation agents, the exponentially large number of probable pose and action variations make the problem challenging and non-trivial. However, the existing datasets and methods for human affordance prediction in 2D scenes are significantly limited in the literature. In this paper, we propose a novel cross-attention mechanism to encode the scene context for affordance prediction by mutually attending spatial feature maps from two different modalities. The proposed method is disentangled among individual subtasks to efficiently reduce the problem complexity. First, we sample a probable location for a person within the scene using a variational autoencoder (VAE) conditioned on the global scene context encoding. Next, we predict a potential pose template from a set of existing human pose candidates using a classifier on the local context encoding around the predicted location. In the subsequent steps, we use two VAEs to sample the scale and deformation parameters for the predicted pose template by conditioning on the local context and template class. Our experiments show significant improvements over the previous baseline of human affordance injection into complex 2D scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。