提出首个音视频鱼类喂食强度增量学习框架,解决新鱼种适应难题。
Audio-Visual Class-Incremental Learning for Fish Feeding intensity Assessment in Aquaculture
- 基于原型的双路径知识保留机制,无须存储历史数据。
- 在6种鱼、81932段音视频数据上实现更高准确率与低存储开销。
- 动态调节音视频权重,适配不同喂食阶段,适合水产养殖智能监控。
鱼类喂食强度评估(FFIA)对工业化水产养殖管理至关重要。现有多模态方法虽提升鲁棒性与效率,但在面对新鱼种或环境时易受灾难性遗忘和数据不足影响。为此,我们首次构建了AV-CIL-FFIA数据集,包含81,932个标注音视频片段,覆盖六种真实养殖环境下的鱼种。进一步提出音视频类增量学习(CIL)框架HAIL-FFIA,采用原型驱动的无样本方法,在不依赖历史数据的前提下,通过分层表示学习与双路径知识保留机制,分离通用喂食强度与鱼种特异性特征。同时引入动态模态平衡系统,自适应调整音频与视觉信息权重。实验表明,该方法在AV-CIL-FFIA上显著优于现有SOTA方法,兼具高精度、低存储需求及强抗遗忘能力。
原文摘要 · Abstract (English)
Fish Feeding Intensity Assessment (FFIA) is crucial in industrial aquaculture management. Recent multi-modal approaches have shown promise in improving FFIA robustness and efficiency. However, these methods face significant challenges when adapting to new fish species or environments due to catastrophic forgetting and the lack of suitable datasets. To address these limitations, we first introduce AV-CIL-FFIA, a new dataset comprising 81,932 labelled audio-visual clips capturing feeding intensities across six different fish species in real aquaculture environments. Then, we pioneer audio-visual class incremental learning (CIL) for FFIA and demonstrate through benchmarking on AV-CIL-FFIA that it significantly outperforms single-modality methods. Existing CIL methods rely heavily on historical data. Exemplar-based approaches store raw samples, creating storage challenges, while exemplar-free methods avoid data storage but struggle to distinguish subtle feeding intensity variations across different fish species. To overcome these limitations, we introduce HAIL-FFIA, a novel audio-visual class-incremental learning framework that bridges this gap with a prototype-based approach that achieves exemplar-free efficiency while preserving essential knowledge through compact feature representations. Specifically, HAIL-FFIA employs hierarchical representation learning with a dual-path knowledge preservation mechanism that separates general intensity knowledge from fish-specific characteristics. Additionally, it features a dynamic modality balancing system that adaptively adjusts the importance of audio versus visual information based on feeding behaviour stages. Experimental results show that HAIL-FFIA is superior to SOTA methods on AV-CIL-FFIA, achieving higher accuracy with lower storage needs while effectively mitigating catastrophic forgetting in incremental fish species learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。