arXiv:2602.09518cs.CV2026-02

构建通用动作空间,统一分析人与灵长类动物行为。

A Universal Action Space for General Behavior Analysis

  • 基于已有动作数据集构建通用动作表征空间。
  • 在多个哺乳动物行为数据集上实现跨物种行为分类。
  • 为跨物种行为分析提供可迁移的视觉表征基础。

动物与人类行为分析一直是计算机视觉中的挑战性任务。早期方法(1970–1990年代)依赖手工设计的边缘检测、分割及颜色、形状、纹理等低层特征来定位物体并推断其身份——这一问题本质不明确。当时的行为分析通常通过跟踪已识别物体并用稀疏特征点建模轨迹,进一步限制了鲁棒性与泛化能力。2010年,Deng与Li提出的ImageNet推动了深度神经网络的大规模视觉识别,使物体识别从复杂的低层处理转向学习高层表征。本文沿此范式,利用现有标注的人类动作数据集构建大规模通用动作空间(Universal Action Space, UAS),并以此为基础分析和分类哺乳动物及黑猩猩行为数据集。代码已开源:https://github.com/franktpmvu/Universal-Action-Space。

原文摘要 · Abstract (English)

Analyzing animal and human behavior has long been a challenging task in computer vision. Early approaches from the 1970s to the 1990s relied on hand-crafted edge detection, segmentation, and low-level features such as color, shape, and texture to locate objects and infer their identities-an inherently ill-posed problem. Behavior analysis in this era typically proceeded by tracking identified objects over time and modeling their trajectories using sparse feature points, which further limited robustness and generalization. A major shift occurred with the introduction of ImageNet by Deng and Li in 2010, which enabled large-scale visual recognition through deep neural networks and effectively served as a comprehensive visual dictionary. This development allowed object recognition to move beyond complex low-level processing toward learned high-level representations. In this work, we follow this paradigm to build a large-scale Universal Action Space (UAS) using existing labeled human-action datasets. We then use this UAS as the foundation for analyzing and categorizing mammalian and chimpanzee behavior datasets. The source code is released on GitHub at https://github.com/franktpmvu/Universal-Action-Space.

行为分析通用表征跨物种动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。