用声音和图像融合判断菠萝保鲜期,准确率超84%。
Classifying Shelf Life Quality of Pineapples by Combining Audio and Visual Features
- 融合多角度音视频特征,构建跨模态分类模型。
- 音频主导采样使准确率达84%,优于单一模态模型。
- 适用于果蔬品质无损检测,适合农业智能化场景。
采用非破坏性方法评估菠萝保鲜质量是减少浪费、提高收益的关键。本文构建了一个多模态、多视角的分类模型,基于声音与视觉特征将菠萝分为四个质量等级。为此,我们整理并发布了包含500个菠萝的PQC500数据集,包含两种模态:通过多个麦克风敲击记录声音,以及从不同位置用多台相机拍摄图像,提供多视角音视频特征。我们改进了对比式音视频掩码自编码器,利用丰富的音视频配对组合训练跨模态分类模型。同时提出采用紧凑采样策略以提升计算效率。在多种数据与模型配置下进行实验,结果表明,使用音频主导采样的跨模态模型达到84%的准确率,分别优于仅用音频和仅用视觉的单模态模型6%和18%。
原文摘要 · Abstract (English)
Determining the shelf life quality of pineapples using non-destructive methods is a crucial step to reduce waste and increase income. In this paper, a multimodal and multiview classification model was constructed to classify pineapples into four quality levels based on audio and visual characteristics. For research purposes, we compiled and released the PQC500 dataset consisting of 500 pineapples with two modalities: one was tapping pineapples to record sounds by multiple microphones and the other was taking pictures by multiple cameras at different locations, providing multimodal and multi-view audiovisual features. We modified the contrastive audiovisual masked autoencoder to train the cross-modal-based classification model by abundant combinations of audio and visual pairs. In addition, we proposed to sample a compact size of training data for efficient computation. The experiments were evaluated under various data and model configurations, and the results demonstrated that the proposed cross-modal model trained using audio-major sampling can yield 84% accuracy, outperforming the unimodal models of only audio and only visual by 6% and 18%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。