arXiv:2503.10212cs.CV2025-03被引 12

用视觉语言模型分析小鼠行为,无需人工标注即可发现新行为模式。

MouseGPT: A Large-scale Vision-Language Model for Mouse Behavior Analysis

  • 构建视觉-语言模型,结合图像与自然语言理解小鼠动作
  • 基于4200万帧数据,实现行为聚类与新行为自动发现
  • 适合神经科学研究者,尤其关注动物行为定量分析的团队

动物行为分析对推动神经科学至关重要,但量化和解读其复杂动态仍具挑战。传统机器视觉方法虽能检测自发行为,却因可解释性差且依赖人工标注,难以覆盖完整行为谱。本文提出MouseGPT,一种视觉-语言模型(VLM),融合视觉线索与自然语言,革新小鼠行为分析。该模型基于首个大规模数据集——包含超过4200万帧、涵盖多种精神疾病状态的姿势动态与开放词汇行为标注——提供上下文丰富的全面行为解析。整体分析框架支持细致的行为画像、聚类与新行为发现,无需耗时的人工标注。评估显示,MouseGPT在精度、适应性与描述丰富性上均超越现有模型,是行为学研究与解析动物模型复杂行为动态的变革性工具。

原文摘要 · Abstract (English)

Analyzing animal behavior is crucial in advancing neuroscience, yet quantifying and deciphering its intricate dynamics remains a significant challenge. Traditional machine vision approaches, despite their ability to detect spontaneous behaviors, fall short due to limited interpretability and reliance on manual labeling, which restricts the exploration of the full behavioral spectrum. Here, we introduce MouseGPT, a Vision-Language Model (VLM) that integrates visual cues with natural language to revolutionize mouse behavior analysis. Built upon our first-of-its-kind dataset - incorporating pose dynamics and open-vocabulary behavioral annotations across over 42 million frames of diverse psychiatric conditions - MouseGPT provides a novel, context-rich method for comprehensive behavior interpretation. Our holistic analysis framework enables detailed behavior profiling, clustering, and novel behavior discovery, offering deep insights without the need for labor - intensive manual annotation. Evaluations reveal that MouseGPT surpasses existing models in precision, adaptability, and descriptive richness, positioning it as a transformative tool for ethology and for unraveling complex behavioral dynamics in animal models.

行为分析视觉语言模型小鼠研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。