arXiv:2501.18543cs.CVcs.RO2025-01被引 1

用视觉Transformer学习人类运动规律,提升轨迹预测精度。

Learning Priors of Human Motion With Vision Transformers

  • 基于视觉Transformer建模人类运动的空间关联性
  • 在标准数据集上优于传统CNN方法
  • 适合城市交通分析与人机共存环境导航

准确理解人类在场景中的移动路径、速度和停顿位置,对城市交通研究和人机共存环境中的机器人导航至关重要。本文提出一种基于视觉变压器(Vision Transformers, ViTs)的神经架构,以提供此类信息。该方案相比卷积神经网络(CNNs)能更有效地捕捉空间相关性。文中详细描述了方法和所提出的神经架构,并展示了在标准数据集上的实验结果。结果显示,所提出的ViT架构在指标上优于基于CNN的方法。

原文摘要 · Abstract (English)

A clear understanding of where humans move in a scenario, their usual paths and speeds, and where they stop, is very important for different applications, such as mobility studies in urban areas or robot navigation tasks within human-populated environments. We propose in this article, a neural architecture based on Vision Transformers (ViTs) to provide this information. This solution can arguably capture spatial correlations more effectively than Convolutional Neural Networks (CNNs). In the paper, we describe the methodology and proposed neural architecture and show the experiments' results with a standard dataset. We show that the proposed ViT architecture improves the metrics compared to a method based on a CNN.

视觉变换器运动预测轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。