通过自适应路由优化矿物图像分类,提升相似类别区分能力。
RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification

- 用动态路由替代固定多尺度聚合,实现样本级尺度偏好选择。
- 在三个数据集上均优于Mona,平均准确率显著提升。
- 适合需要高精度区分相似矿物的地质图像分析场景。
矿物图像分类对地质勘探和资源开发至关重要,但因类内外观差异大、类间视觉相似度高而极具挑战。多认知视觉适配器(Mona)是一种参数高效适配器,仅微调少量参数即可适配预训练视觉模型。然而,Mona采用静态多尺度聚合,难以适应样本特有的尺度偏好,且无法有效缓解视觉相似矿物类别间的混淆问题。为此,我们提出轻量级路由空间正则化方法 RouteGraph-Mona。具体地,将 Mona 的静态多尺度聚合替换为样本自适应路由,形成的分支选择行为构建出紧凑的路由空间,捕捉每张图像的尺度偏好。随后,利用类别专属路由锚点和混淆加权边界对路由签名进行正则化:路由锚点促使类别一致的路由模式,边界增强视觉相似类别在路由空间中的分离性。在两个视觉骨干网络和三个公开矿物图像数据集上的实验表明,RouteGraph-Mona 在平均准确率上持续优于 Mona,且与主流微调方法及矿物图像分类基线保持竞争力。
原文摘要 · Abstract (English)
Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-class variations in appearance and high inter-class visual similarity. Multi-cognitive Visual Adapter (Mona) is a vision-oriented parameter-efficient adapter that adapts pre-trained visual models by tuning only a few parameters. However, Mona statically aggregates responses from multiple scales, limiting its ability to accommodate sample-specific scale preferences and model confusion among visually similar mineral categories. To address this issue, we propose \textbf{RouteGraph-Mona}, a lightweight route-space regularization method built on Mona. Specifically, we replace Mona's static multi-scale aggregation with sample-adaptive routing. The resulting branch-selection behavior defines a compact routing space that captures each image's scale preferences. We then regularize the resulting routing signatures with class-wise route anchors and confusion-weighted margins. The route anchors encourage class-consistent routing patterns, while the margins promote greater separation between visually similar categories in the routing space. Experiments on three public mineral image datasets with two visual backbones show that RouteGraph-Mona consistently outperforms Mona in mean accuracy and remains competitive with representative fine-tuning methods and mineral image classification baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。