评测大模型对宠物狗视频的理解能力,发现现有模型在长期互动推理上表现有限。
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

- 基于网络视频自动生成5000条狗行为问答数据,支持多跳推理
- 907段真实家庭狗视频覆盖5类任务,测试长时序理解能力
- 提出可复用的数据构建方法,适用于低数据领域
多模态大模型在图像、视频、音频和文本等多样化输入上展现出强大的零样本能力。然而,其在动物为中心场景中的应用仍鲜有研究。鉴于动物在千万家庭中的重要性,对宠物相关任务(如识别应激信号、实现响应式机器人陪伴)进行基准评测,对构建能与人类协同的下一代AI系统至关重要。本文提出K9-Bench,一个专注于真实家庭狗视频的新基准,包含约5000个问答对,覆盖907段视频及5个任务类别,旨在测试多模态大模型在狗行为与交互理解上的长时序多模态推理能力。我们设计了一套可扩展的、基于视觉语言模型(VLM)与大语言模型(LLM)的数据生成管道,自动从网络挖掘狗相关视频并精炼需要细粒度、多跳推理的问答对。通过引入偏见缓解策略,有效减少由VLM在数据筛选中引入的偏差。大量实验表明,前沿多模态大模型在狗行为任务上的零样本表现有限:尽管闭源模型优于开源模型,但对分散在长时序中的细微姿态与互动线索仍难以进行组合推理。通用链式思维提示仅带来有限性能提升。除提供狗活动分析新数据集外,K9-Bench还提供可迁移的数据构建框架,适用于其他低数据领域量化评估。项目主页:https://ogmenrobotics.github.io/K9Bench。
原文摘要 · Abstract (English)
MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies in understanding and modeling animal-centric scenarios. As animals are integral to millions of households, benchmarking next-generation AI models on pet-focused tasks, ranging from recognizing distress signals to enabling responsive robotic companions, is essential for building AI systems that can work alongside us. We introduce K9-Bench, a novel benchmark focused on real-world domestic dog videos, specifically targeting canine action and interaction understanding via approximately 5000 question-answer pairs across 907 videos spanning 5 distinct task categories that test long-form, canine-centric multimodal reasoning in MLLMs. To create this dataset, we propose a scalable, VLM/LLM-powered data generation pipeline that automatically mines canine-centric videos from the web and curates QA pairs requiring fine-grained, multi-hop reasoning over canine actions and temporally extended interaction sequences. We implement bias mitigation strategies designed to eliminate biases introduced by VLMs during dataset curation. Through extensive experimentation, we find that frontier MLLMs exhibit limited zero-shot performance on canine-centric tasks: although state-of-the-art closed-source models outperform open-source counterparts, they still struggle with compositional reasoning over subtle posture and interaction cues spread over long horizons. We observe that generic chain-of-thought prompting provides only modest performance for such long-horizon reasoning. Beyond a novel dataset for canine activity analysis, K9-Bench provides a general-purpose dataset construction pipeline that can be adapted to other low-data domains for quantitative analysis. Our project website is available at: https://ogmenrobotics.github.io/K9Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。