用多视角视频自监督学习动物3D行为,无需人工标注
BEAST3D: Animal behavioral analysis and neural encoding from multi-view video via Gaussian splatting

- 通过视觉变换器预测3D高斯点云,实现无标注多视角重建
- 仅需4个视角即可还原3D结构,准确率在四物种上验证
- 适用于行为分析、神经编码等场景,适合实验室动物研究
多视角视频在实验环境中越来越广泛用于捕捉动物的3D运动,但从中提取丰富的3D表征仍具挑战。监督姿态估计依赖大量人工标注,而通用3D重建模型在实验室特有的稀疏视角和特殊成像条件下表现不佳。我们提出BEAST3D,一种自监督预训练框架,从未标注的标定多视角视频中学习3D视觉表征。BEAST3D使用视觉变换器预测3D高斯点云,通过可微渲染重建被遮挡视角,同时分割动物与背景。该方法直接利用已知相机参数,仅需4个视角即可重建3D结构——不同于通用模型需依赖密集重叠视角估算相机位姿,这在实验室中极少具备。在四种物种上的综合评估表明,BEAST3D生成的丰富且视角不变的特征可有效迁移至三项下游任务:新视角合成(验证3D表征质量)、多视角姿态估计(提供行为分析常用的关键点轨迹)以及神经编码(关联3D行为特征与同步记录的神经活动)。因此,BEAST3D为现代多视角实验记录中的行为分析建立了一个灵活高效的3D结构利用框架。
原文摘要 · Abstract (English)
Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representations from these recordings remains challenging. Supervised pose estimation requires extensive manual annotation, while general-purpose 3D reconstruction models trained on generic scene datasets fail on the specialized imagery and sparse-view setting of laboratory experiments. We address these limitations with BEAST3D, a self-supervised pretraining framework that learns 3D visual representations from unlabeled, calibrated multi-view video. BEAST3D uses a vision transformer to predict 3D Gaussian splats that reconstruct held-out views through differentiable rendering, while simultaneously segmenting the animal from the background. BEAST3D reconstructs 3D structure with as few as four views by conditioning directly on known camera parameters--unlike general-purpose models, which must estimate camera geometry from dense overlapping viewpoints that are seldom available in lab settings. Through comprehensive evaluation across four species, we demonstrate that BEAST3D produces rich, viewpoint-invariant features that transfer effectively to three downstream tasks: novel view synthesis, which validates the quality of the learned 3D representations; multi-view pose estimation, which provides the sparse keypoint trajectories widely used in behavioral analysis; and neural encoding, which relates 3D behavioral features to simultaneously recorded neural activity. BEAST3D thus establishes a versatile framework for behavioral analysis that leverages 3D structure in modern multi-view laboratory recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。