系统梳理多模态动物姿态估计方法与数据集,助力生物医学研究。
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
- 按传感器类型分类176篇论文,分析多模态输入方法与学习范式。
- 总结2D/3D动物姿态数据集与评估指标,涵盖视觉、红外、惯性等模态。
- 揭示动物与人体姿态估计的相互促进关系,适合跨学科研究者参考。
动物姿态估计(APE)旨在利用多种传感器和模态输入(如RGB相机、激光雷达、红外、惯性测量单元、声学及语言线索)定位动物身体部位,对神经科学、生物力学和兽医学研究至关重要。通过对2011年以来176篇论文的评估,本文按输入传感器与模态类型、输出形式、学习范式、实验设置及应用领域对APE方法进行分类,并深入分析单模态与多模态系统当前趋势、挑战与未来方向。研究还探讨了人与动物姿态估计之间的相互转化关系,指出APE创新如何反哺人体姿态估计及更广泛的机器学习范式。此外,文中提供了基于不同传感器与模态的2D与3D APE数据集及评估指标。项目主页持续更新:https://github.com/ChennyDeng/MM-APE。
原文摘要 · Abstract (English)
Animal pose estimation (APE) aims to locate the animal body parts using a diverse array of sensor and modality inputs (e.g. RGB cameras, LiDAR, infrared, IMU, acoustic and language cues), which is crucial for research across neuroscience, biomechanics, and veterinary medicine. By evaluating 176 papers since 2011, APE methods are categorised by their input sensor and modality types, output forms, learning paradigms, experimental setup, and application domains, presenting detailed analyses of current trends, challenges, and future directions in single- and multi-modality APE systems. The analysis also highlights the transition between human and animal pose estimation, and how innovations in APE can reciprocally enrich human pose estimation and the broader machine learning paradigm. Additionally, 2D and 3D APE datasets and evaluation metrics based on different sensors and modalities are provided. A regularly updated project page is provided here: https://github.com/ChennyDeng/MM-APE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。