无需微调,用视觉语言模型实现动物行为的无监督姿态与行为理解。
BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning
- 通过多阶段推理框架,利用量子点数据减少人工标注。
- 姿态估计中引入重投影误差检测低置信度标签,提升可靠性。
- 直接从视频分析行为,无需关键点,适合神经科学多动物研究。
理解自由运动动物行为是神经科学的核心,姿态估计与行为理解是连接神经活动与自然动作的基础。然而,这两项任务仍严重依赖人工标注或不稳定的无监督流程,限制了可扩展性与可重复性。我们提出 BehaviorVLM,一个统一的视觉-语言框架,可在无需特定任务微调和极少人工标注的情况下,通过详细、明确且可验证的推理步骤引导预训练视觉-语言模型(VLM)。在姿态估计方面,我们结合量子点标记的行为数据,设计多阶段流程,融合时空与跨视角推理,显著降低人工标注负担,通过重投影误差等几何检验暴露低置信度标签,并生成可后期过滤、修正或用于微调下游姿态模型的标签。在行为理解方面,提出整合深度嵌入聚类以发现过分割行为片段、基于VLM的每片段视频描述生成,以及基于LLM的推理来合并并语义标注行为段。该行为分析流程可直接从视觉信息出发,无需关键点即可分割行为。整体框架实现了可扩展、可解释且标签轻量的多动物行为分析。
原文摘要 · Abstract (English)
Understanding freely moving animal behavior is central to neuroscience, where pose estimation and behavioral understanding form the foundation for linking neural activity to natural actions. Yet both tasks still depend heavily on human annotation or unstable unsupervised pipelines, limiting scalability and reproducibility. We present BehaviorVLM, a unified vision-language framework for pose estimation and behavioral understanding that requires no task-specific finetuning and minimal human labeling by guiding pretrained Vision-Language Models (VLMs) through detailed, explicit, and verifiable reasoning steps. For pose estimation, we leverage quantum-dot-grounded behavioral data and propose a multi-stage pipeline that integrates temporal, spatial, and cross-view reasoning. This design greatly reduces human annotation effort, exposes low-confidence labels through geometric checks such as reprojection error, and produces labels that can later be filtered, corrected, or used to fine-tune downstream pose models. For behavioral understanding, we propose a pipeline that integrates deep embedded clustering for over-segmented behavior discovery, VLM-based per-clip video captioning, and LLM-based reasoning to merge and semantically label behavioral segments. The behavioral pipeline can operate directly from visual information and does not require keypoints to segment behavior. Together, these components enable scalable, interpretable, and label-light analysis of multi-animal behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。