评测大模型标注老鼠行为视频的能力,发现现有模型表现不佳。
Rodent-Bench
- 构建多模态大模型行为评估基准,涵盖社交、梳理等多元行为。
- 模型在长视频中难以准确分割时间片段,整体性能普遍偏低。
- 适合关注神经科学自动化标注的科研人员参考。
我们提出Rodent-Bench,一个用于评估多模态大语言模型(MLLMs)标注老鼠行为视频能力的新基准。该基准覆盖多种行为范式,包括社交互动、梳理、抓挠和冻结行为,视频时长介于10至35分钟。我们评估了Gemini-2.5-Pro、Gemini-2.5-Flash和Qwen-VL-Max等先进MLLMs,发现它们均未达到可实用水平。基准提供两种版本以适应不同模型能力,并设立标准化评估指标,包括逐秒准确率、宏平均F1、平均精度均值、互信息和马修斯相关系数。尽管部分模型在梳理检测任务上表现尚可,但整体在时间分段、长视频处理及细微行为区分方面仍存在显著挑战。分析揭示了当前MLLMs在科学视频标注中的关键局限,为未来模型改进提供方向。Rodent-Bench为推动神经科学研究中可靠自动化行为标注的发展奠定基础。
原文摘要 · Abstract (English)
We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-the-art MLLMs, including Gemini-2.5-Pro, Gemini-2.5-Flash and Qwen-VL-Max, using this benchmark and find that none of these models perform strongly enough to be used as an assistant for this task. Our benchmark encompasses diverse datasets spanning multiple behavioral paradigms including social interactions, grooming, scratching, and freezing behaviors, with videos ranging from 10 minutes to 35 minutes in length. We provide two benchmark versions to accommodate varying model capabilities and establish standardized evaluation metrics including second-wise accuracy, macro F1, mean average precision, mutual information, and Matthew's correlation coefficient. While some models show modest performance on certain datasets (notably grooming detection), overall results reveal significant challenges in temporal segmentation, handling extended video sequences, and distinguishing subtle behavioral states. Our analysis identifies key limitations in current MLLMs for scientific video annotation and provides insights for future model development. Rodent-Bench serves as a foundation for tracking progress toward reliable automated behavioral annotation in neuroscience research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。