arXiv:2605.07903cs.SDcs.AI2026-05

无需标注,从蜂群嗡鸣声中自动发现可重复的声学状态。

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

论文配图:BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing
图 1 · 摘自论文原文
  • 用自监督模型提取音频特征,再用无监督编码器学习离散声学标记。
  • 标记能区分有王与无王蜂群,且无王状态可进一步分解为三个稳定子状态。
  • 结果在不同实验条件下一致,适合研究蜂群行为或开发蜂巢健康监测系统。

在无监督条件下发现生物信号中的结构是计算智能的核心问题,但现有生物声学方法依赖发声模型或预定义语义单元,难以适用于非发声物种。本文提出BeeVe,一种针对集体蜜蜂嗡鸣声的无监督声学状态发现框架。BeeVe采用冻结的自监督Patchout Spectrogram Transformer(PaSST)作为特征提取器,对提取的嵌入向量训练无标签的向量量化变分自编码器(VQ-VAE),直接从未标注蜂巢音频中学习有限离散的声学标记。整个过程不使用标签、预训练任务或对比目标。后验评估显示,学习到的标记在已知蜂王状态下的杰恩-申诺尔散度值介于0.609至0.688之间,能有效区分有王与无王状态;且无王状态可进一步分解为三个内部一致的子状态,在不同码本大小和随机种子下均保持稳定。标记转移分析表明所有实验中存在非随机的序列结构(p << 0.001)。对未见录音的泛化测试显示,标记重叠率(Jaccard = 0.947)和全局流形拓扑结构得以保留。结果表明,无监督离散码本学习可在无需标注的情况下从非发声生物信号中恢复可重复的声学结构,为无创蜂巢健康监测提供新路径。

原文摘要 · Abstract (English)

Discovering structure in biological signals without supervision is a fundamental problem in computational intelligence, yet existing bioacoustic methods assume vocal production models or predefined semantic units, leaving non-vocal species poorly served. This work introduces BeeVe, an unsupervised framework for acoustic state discovery in collective honey bee buzzing. BeeVe uses the self-supervised Patchout Spectrogram Transformer (PaSST) as a frozen feature extractor, then trains a Vector-Quantized Variational Autoencoder (VQ-VAE) without labels on those embeddings, learning a finite discrete codebook of acoustic tokens directly from unlabelled hive audio. No labels, pretext tasks, or contrastive objectives are used at any stage. Post-hoc evaluation against known queen status reveals that the learned tokens separate queenright and queenless conditions with Jensen-Shannon Divergence values between 0.609 and 0.688, and that the queenless condition further decomposes into three internally coherent sub-states stable across experiments with different codebook sizes and random seeds. Token transition analysis confirms non-random sequential structure (p << 0.001) across all experiments. Generalisation to unseen recordings preserves both token overlap (Jaccard = 0.947) and global manifold topology. These results demonstrate that unsupervised discrete codebook learning can recover repeatable acoustic structure from a non-vocal biological signal without annotation, opening a path toward non-invasive acoustic hive health monitoring.

声学分析无监督学习蜂群行为生物信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。