arXiv:2509.05703cs.CVcs.AI2025-09

用视觉语言模型分析海洋哺乳动物声谱图,无需标注就能自动理解声音模式。

Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis

  • 将视觉语言模型与大模型验证结合,实现声谱图的自动解读。
  • 在无手动标注情况下,成功识别出多种海洋哺乳动物的发声特征。
  • 适合生物声学研究者、海洋保护项目和自动化监测系统使用。

海洋哺乳动物的发声分析依赖于对生物声学频谱图的解读。视觉语言模型(VLMs)并未在这些领域特定的可视化数据上进行训练。本文研究了VLMs是否能从频谱图中提取有意义的模式。所提出的框架将VLM的视觉解释能力与基于大语言模型(LLM)的验证机制相结合,构建领域知识体系。该方法可实现对声学数据的适应性分析,无需人工标注或模型重新训练,显著降低分析成本并提升自动化水平。

原文摘要 · Abstract (English)

Marine mammal vocalization analysis depends on interpreting bioacoustic spectrograms. Vision Language Models (VLMs) are not trained on these domain-specific visualizations. We investigate whether VLMs can extract meaningful patterns from spectrograms visually. Our framework integrates VLM interpretation with LLM-based validation to build domain knowledge. This enables adaptation to acoustic data without manual annotation or model retraining.

声谱分析视觉语言模型海洋生物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。