arXiv:2609.08570cs.CV2026-09

对比深度学习与传统图像分析,发现模型架构和采样策略显著影响污泥显微图像识别效果。

Effects of model architecture and learning strategies on deep learning-based recognition of activated sludge microscopic images and comparison with quantitative image analysis

论文配图:Effects of model architecture and learning strategies on deep learning-based recognition of activated sludge microscopic images and comparison with quantitative image analysis
图 1 · 摘自论文原文
  • 采用Transformer和自监督预训练提升分类准确率
  • 过度下采样降低精度,保持视野比保持分辨率更有效
  • 深度学习显著优于传统定量图像分析方法

显微图像分析长期以来被视为监测活性污泥的有前景方法。近年来,由于表现优异,基于深度学习的图像分析在该领域日益普及。然而,以往研究多依赖卷积神经网络(CNN)和ImageNet监督预训练,较少探索基于Transformer的模型或自监督基础模型。此外,多数研究对图像进行下采样,但下采样策略的影响尚未充分研究,其与分析性能的关系仍不明确。更关键的是,尚无研究对深度学习性能与此前广泛使用的定量图像分析(QIA)进行定量比较。本研究准备了三种活性污泥样本,对显微图像进行分类并评估分类准确率。结果表明,基于Transformer的架构和替代预训练方法在分类准确率上表现更优。下采样分析显示,图像过小会降低准确率,但超过一定尺寸后继续增大图像无法进一步提升性能。同时,保持视场范围比保持分辨率更有利于提高分类准确率。最后,深度学习在准确率上显著优于定量图像分析。

原文摘要 · Abstract (English)

Microscopic image analysis has long been recognized as a promising approach for monitoring activated sludge. In recent years, deep learning-based image analysis has been increasingly adopted in this field because of its high performance. However, previous studies on microscopic image analysis of activated sludge have rarely explored transformer-based models or self-supervised foundation models and have instead relied on CNNs and supervised ImageNet pretraining. In addition, previous studies often downsampled image sizes, but the effects of downsampling have not been sufficiently investigated, and the relationship between downsampling strategies and image analysis performance remains unclear. Furthermore, no study has quantitatively compared deep learning performance with quantitative image analysis (QIA), which was widely used before the emergence of deep learning. In this study, to examine how model architecture and learning strategies affect performance in microscopic image analysis of activated sludge and to quantitatively determine whether deep learning outperforms QIA, we prepared three types of activated sludge samples, classified their microscopic images, and evaluated classification accuracy. Our results showed that transformer-based architectures and alternative pretraining methods were effective in terms of classification accuracy. Our downsampling analysis showed that using overly small images reduced accuracy, but increasing image size beyond a certain point did not improve it further. In addition, the analysis indicated that, to achieve high classification accuracy, maintaining the field of view was a more effective downsampling strategy than maintaining resolution. Finally, our comparison between deep learning and QIA showed that deep learning outperformed QIA in terms of accuracy.

图像识别深度学习活性污泥量化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。