arXiv:2507.09513q-bio.NCcs.CV2025-07被引 6

用自监督预训练提升动物行为分析精度,少标注也能用。

Animal behavioral analysis and neural encoding with transformer-based self-supervised pretraining

  • 用Transformer结合掩码重建和时序对比学习,从无标签视频中学特征。
  • 在多物种、单/多动物场景下,行为特征提取、姿态估计和动作分割均更准。
  • 适合标注数据少的神经行为研究,可直接当骨干模型用。

大脑只有通过其产生的行为才能被完整理解——这是现代神经科学的核心原则,但带来显著技术挑战。许多研究用摄像头捕捉行为,但传统视频分析依赖需大量标注数据的专用模型。我们提出BEAST(基于Transformer的自监督预训练行为分析),一种新型可扩展框架,针对特定实验预训练视觉Transformer,适用于多种神经行为分析任务。BEAST结合掩码自编码与时序对比学习,有效利用无标签视频数据。在多物种上的综合评估显示,该方法在三项关键神经行为任务中表现更优:提取与神经活动相关的行为特征,以及单/多动物场景下的姿态估计与动作分割。本方法建立了一个强大且通用的主干模型,显著加速标注数据稀缺场景下的行为分析。

原文摘要 · Abstract (English)

The brain can only be fully understood through the lens of the behavior it generates -- a guiding principle in modern neuroscience research that nevertheless presents significant technical challenges. Many studies capture behavior with cameras, but video analysis approaches typically rely on specialized models requiring extensive labeled data. We address this limitation with BEAST(BEhavioral Analysis via Self-supervised pretraining of Transformers), a novel and scalable framework that pretrains experiment-specific vision transformers for diverse neuro-behavior analyses. BEAST combines masked autoencoding with temporal contrastive learning to effectively leverage unlabeled video data. Through comprehensive evaluation across multiple species, we demonstrate improved performance in three critical neuro-behavioral tasks: extracting behavioral features that correlate with neural activity, and pose estimation and action segmentation in both the single- and multi-animal settings. Our method establishes a powerful and versatile backbone model that accelerates behavioral analysis in scenarios where labeled data remains scarce.

行为分析自监督视觉模型神经科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。