arXiv:2411.01455cs.CVcs.MA2024-11被引 2

用分层记忆增强的Transformer模型预测多人互动中的未来动作。

HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation

  • 分层记忆架构融合全局历史信息与局部特征。
  • 在多场景测试中显著优于现有最先进方法。
  • 适合需要理解群体行为的机器人与自动驾驶研究者。

理解与预测人类行为是机器人智能感知的核心挑战。尽管个体行为预测已取得进展,但真实场景中的人类活动往往涉及交互,而现有工作对此关注不足。为此,本文提出分层记忆感知变换器(HiMemFormer),一种用于在线多智能体动作预测的Transformer模型。该模型通过变换器框架整合并分发跨所有智能体的全局历史信息,采用从粗到细的策略,由分层本地记忆解码器基于全局表示解析每个智能体的特定特征。相比以往方法,HiMemFormer通过层次化地将全局上下文与个体偏好结合,有效避免了多智能体预测中的噪声与冗余信息。在多种多智能体场景下的大量实验表明,该模型性能显著优于其他先进方法。

原文摘要 · Abstract (English)

Understanding and predicting human actions has been a long-standing challenge and is a crucial measure of perception in robotics AI. While significant progress has been made in anticipating the future actions of individual agents, prior work has largely overlooked a key aspect of real-world human activity -- interactions. To address this gap in human-like forecasting within multi-agent environments, we present the Hierarchical Memory-Aware Transformer (HiMemFormer), a transformer-based model for online multi-agent action anticipation. HiMemFormer integrates and distributes global memory that captures joint historical information across all agents through a transformer framework, with a hierarchical local memory decoder that interprets agent-specific features based on these global representations using a coarse-to-fine strategy. In contrast to previous approaches, HiMemFormer uniquely hierarchically applies the global context with agent-specific preferences to avoid noisy or redundant information in multi-agent action anticipation. Extensive experiments on various multi-agent scenarios demonstrate the significant performance of HiMemFormer, compared with other state-of-the-art methods.

多智能体动作预测Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。