arXiv:2603.07294cs.CVcs.AI2026-03EMNLP被引 3

专为鸟类物种设计的多模态对话助手,提升生态研究信息获取效率。

MAviS: A Multimodal Conversational Assistant For Avian Species

  • 构建涵盖1000+鸟种的多模态数据集,整合图像、音频与文本
  • 在2.5万+问答对上测试,性能显著超越基线模型
  • 适合生态保护、生态监测等领域的研究人员使用

细粒度的物种理解与特定物种的多模态问答对推进生物多样性保护和生态监测至关重要。然而,现有多模态大模型在鸟类等专业领域面临挑战,难以提供准确且情境相关的信息。为此,我们提出MAviS-Dataset,一个大规模多模态鸟类数据集,覆盖超过1000种鸟类,包含图像、音频与文本模态,并配有预训练与指令微调子集,富含结构化问答对。基于该数据集,我们开发了MAviS-Chat,一种支持音频、视觉与文本的多模态大模型,专用于精细物种理解、多模态问答及场景描述生成。此外,我们构建了MAviS-Bench,一个包含超25,000个问答对的基准,用于评估跨模态的鸟类感知与推理能力。实验表明,MAviS-Chat在性能上远超基线MiniCPM-o-2.6,达到开源模型领先水平,验证了指令微调的MAviS-Dataset的有效性。研究强调了领域自适应多模态大模型在生态应用中的必要性。

原文摘要 · Abstract (English)

Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models face challenges when it comes to specialized topics like avian species, making it harder to provide accurate and contextually relevant information in these areas. To address this limitation, we introduce the MAviS-Dataset, a large-scale multimodal avian species dataset that integrates image, audio, and text modalities for over 1,000 bird species, comprising both pretraining and instruction-tuning subsets enriched with structured question-answer pairs. Building on the MAviS-Dataset, we introduce MAviS-Chat, a multimodal LLM that supports audio, vision, and text and is designed for fine-grained species understanding, multimodal question answering, and scene-specific description generation. Finally, for quantitative evaluation, we present MAviS-Bench, a benchmark of over 25,000 QA pairs designed to assess avian species-specific perceptual and reasoning abilities across modalities. Experimental results show that MAviS-Chat outperforms the baseline MiniCPM-o-2.6 by a large margin, achieving state-of-the-art open-source results and demonstrating the effectiveness of our instruction-tuned MAviS-Dataset. Our findings highlight the necessity of domain-adaptive multimodal LLMs for ecological applications.

多模态鸟类识别生态智能对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。