arXiv:2510.19160cs.LG2025-10被引 1

用视觉语言模型自动识别小鼠恐惧行为,无需微调即可高精度分类。

Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression

  • 基于Qwen2.5-VL模型,结合提示工程与帧级预处理提升行为识别能力。
  • 在不微调模型前提下,对冻结、逃跑等稀有行为均实现高F1分数。
  • 适合神经科学、行为学等跨学科研究者快速构建多时点行为数据集。

多源数据融合是推动多领域科学探索的关键步骤。本文构建了一个视觉-语言模型(VLM),通过文本输入编码视频,以分类小鼠在环境中表现出的各类行为。该模型可为每个受试小鼠和每次实验生成随时间变化的行为向量,输出高质量数据集,且准确率高、人工干预少。我们采用开源Qwen2.5-VL模型,通过提示工程、带标注样本的上下文学习(ICL)及帧级预处理增强性能。结果显示,各项方法均提升分类效果,联合使用后在所有行为类别上均获得优异F1分数,包括冻结、逃跑等稀有行为,且无需模型微调。该模型将助力跨学科研究人员整合多时间点、多环境下的行为特征,形成综合性数据集,支持复杂科研问题的解答。

原文摘要 · Abstract (English)

Integration of diverse data will be a pivotal step towards improving scientific explorations in many disciplines. This work establishes a vision-language model (VLM) that encodes videos with text input in order to classify various behaviors of a mouse existing in and engaging with their environment. Importantly, this model produces a behavioral vector over time for each subject and for each session the subject undergoes. The output is a valuable dataset that few programs are able to produce with as high accuracy and with minimal user input. Specifically, we use the open-source Qwen2.5-VL model and enhance its performance through prompts, in-context learning (ICL) with labeled examples, and frame-level preprocessing. We found that each of these methods contributes to improved classification, and that combining them results in strong F1 scores across all behaviors, including rare classes like freezing and fleeing, without any model fine-tuning. Overall, this model will support interdisciplinary researchers studying mouse behavior by enabling them to integrate diverse behavioral features, measured across multiple time points and environments, into a comprehensive dataset that can address complex research questions.

行为识别视觉语言模型小鼠实验自动化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。