用多模态大模型分析小鼠行为视频,预测社会等级
MTT-Bench: Predicting Social Dominance in Mice via Multimodal Large Language Models

- 用多模态大模型直接解析原始小鼠行为视频
- 在未标注数据上零样本预测,与管测排名高度一致
- 为动物行为学提供无需定制模型的新方法
理解动物行为中的社会支配关系对神经科学和行为研究至关重要。本文探索多模态大语言模型(MLLMs)分析小鼠原始行为视频并预测其社会等级层次的能力。我们提出MTT-Bench,一个包含成对小鼠互动标注视频的新型基准数据集,用于小鼠管测(Mouse Tube Test)分析。基于现有MLLM架构,我们微调模型以在未见行为序列上进行零样本推理,测试时无需显式标签即可预测社会支配关系。实验结果表明,该框架与管测排名具有高度一致性。本工作为将基础模型应用于动物行为学和社交行为分析开辟了新方向,且无需设计领域专用模型。
原文摘要 · Abstract (English)
Understanding social dominance in animal behavior is critical for neuroscience and behavioral studies. In this work, we explore the capability of Multimodal Large Language Models(MLLMs) to analyze raw behavioral video of mice and predict their dominance hierarchy. We introduce MTT-Bench, a novel benchmark comprising annotated videos of pairwise mouse interactions for Mouse Tube Test analysis. Building on existing MLLM architectures, we fine-tune these models to perform zero-shot inference on unseen behavioral sequences, predicting social dominance without explicit labels during testing. Our framework demonstrates promising results, showing high agreement with tube test rankings. This work opens a new direction for applying foundation models to ethology and social behavior analysis, without the need to design domain-specific models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。