arXiv:2605.09152cs.CLq-bio.NC2026-05被引 1

首个能理解猫行为意图的多模态大模型,结合生理数据精准识别猫的内心状态。

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

论文配图:Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
图 1 · 摘自论文原文
  • 融合视频、音频、生理信号与文本,实现跨模态深层推理。
  • 在新基准上达到71.16%准确率,显著超越现有模型。
  • 开源完整工具链,适合动物行为研究与兽医诊断应用。

解析动物意图是计算行为学的核心挑战,主要源于语义混淆现象——同一外部信号(如猫的呼噜声)在不同生理情境下可能对应截然不同的内在状态。现有多模态大语言模型无法处理高频生物时间序列数据,仅能进行表层行为匹配,难以实现真正的潜在状态推理。为此,我们提出Meow-Omni 1,首个面向计算行为学的开源四模态多模态大模型。它原生融合视频、音频、生理时间序列流与文本推理能力。通过针对性架构改进,将专用科学编码器集成至统一主干网络,并基于生理学基础建立跨模态对齐机制以实现意图推断。在全新专家验证的四模态基准MeowBench上,Meow-Omni 1达到71.16%的意图识别准确率,显著优于领先的视觉-语言及全模态基线模型。我们开源了包含模型权重、训练框架与Meow-10K数据集的完整流程,旨在建立跨物种意图理解的可扩展范式,推动基础模型向真实世界兽医诊断与野生动物保护迈进。

原文摘要 · Abstract (English)

Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical external signals (e.g., a cat's purr) correspond to radically different internal states depending on physiological context. Existing Multimodal Large Language Models (MLLMs) are blind to high-frequency biological time-series data, restricting them to superficial behavioural pattern matching rather than genuine latent-state reasoning. To bridge this gap, we introduce Meow-Omni 1, the first open-source, quad-modal MLLM purpose-built for computational ethology. It natively fuses video, audio, and physiological time-series streams with textual reasoning. Through targeted architectural adaptation, we integrate specialized scientific encoders into a unified backbone and formalize intent inference via physiologically grounded cross-modal alignment. Evaluated on MeowBench, a novel, expert-verified quad-modal benchmark, Meow-Omni 1 achieves state-of-the-art intent-recognition accuracy (71.16%), substantially outperforming leading vision-language and omni-modal baselines. We release the complete open-source pipeline including model weights, training framework, and the Meow-10K dataset, to establish a scalable paradigm for inter-species intent understanding and to advance foundation models toward real-world veterinary diagnostics and wildlife conservation.

多模态动物行为大模型生理信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。