arXiv:2506.12712cs.CVeess.IV2025-06

用新型网络提升煤岩显微组分分割精度,同时大幅降低参数量。

Combining Self-attention and Dilation Convolutional for Semantic Segmentation of Coal Maceral Groups

  • 提出DA-VIT模型,融合自注意力与空洞卷积,优化特征提取。
  • 在煤岩图像上达到92.14%像素准确率和63.18% mIoU,优于现有方法。
  • 参数量仅4.95M,减少81.18%计算开销,适合资源受限场景。

煤岩显微组分的分割可视为煤岩图像的语义分割任务,对研究煤的化学性质具有重要意义。现有模型多通过堆叠参数提升精度,导致计算量大、训练效率低。同时,煤岩图像采样专业性强、样本获取耗时且依赖人工。为此,本文创新性地提出基于物联网的DA-VIT并行网络模型,通过物联网持续扩展数据集,实现分割精度的持续提升。此外,将并行网络与主干网络解耦,保障主干网络在数据更新时正常运行。引入DCSA机制增强微观图像局部特征,将大核卷积注意力分解为多尺度,参数量减少81.18%。对比实验与消融实验表明,DA-VIT-Base在多项指标上表现优异,像素准确率达92.14%,mIoU达63.18%;DA-VIT-Tiny参数量为4.95M,FLOPs为8.99G,各项性能均超越当前先进方法。

原文摘要 · Abstract (English)

The segmentation of coal maceral groups can be described as a semantic segmentation process of coal maceral group images, which is of great significance for studying the chemical properties of coal. Generally, existing semantic segmentation models of coal maceral groups use the method of stacking parameters to achieve higher accuracy. It leads to increased computational requirements and impacts model training efficiency. At the same time, due to the professionalism and diversity of coal maceral group images sampling, obtaining the number of samples for model training requires a long time and professional personnel operation. To address these issues, We have innovatively developed an IoT-based DA-VIT parallel network model. By utilizing this model, we can continuously broaden the dataset through IoT and achieving sustained improvement in the accuracy of coal maceral groups segmentation. Besides, we decouple the parallel network from the backbone network to ensure the normal using of the backbone network during model data updates. Secondly, DCSA mechanism of DA-VIT is introduced to enhance the local feature information of coal microscopic images. This DCSA can decompose the large kernels of convolutional attention into multiple scales and reduce 81.18% of parameters.Finally, we performed the contrast experiment and ablation experiment between DA-VIT and state-of-the-art methods at lots of evaluation metrics. Experimental results show that DA-VIT-Base achieves 92.14% pixel accuracy and 63.18% mIoU. Params and FLOPs of DA-VIT-Tiny are 4.95M and 8.99G, respectively. All of the evaluation metrics of the proposed DA-VIT are better than other state-of-the-art methods.

语义分割煤岩分析轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。