用小模型驱动大模型,动态激活专家提升缺陷检测精度
Distilled Large Language Model-Driven Dynamic Sparse Expert Activation Mechanism
- 用文本引导的稀疏专家混合机制,按语义自动选专家
- 在3个缺陷数据集上mAP提升最高达13.9个百分点
- 适合工业质检场景,兼顾速度与精度
高类别相似性、极端尺度变化以及有限计算资源限制了复杂真实数据中可靠视觉识别。现有视觉主导或跨模态方法常依赖固定融合机制和繁重标注流程,导致泛化性能不佳。我们提出基于蒸馏大语言模型驱动的稀疏专家混合(DS-MoE)框架,融合文本引导的动态路由与轻量级多尺度理解能力。该框架通过稀疏MoE结构,将文本语义与缺陷特定视觉模式动态对齐,依据语义相关性自适应激活任务相关专家,缓解类别间混淆问题。轻量级MobileSAM编码器支持实时推理,同时保持多尺度缺陷细节。在PCB、铝箔和模具缺陷数据集上的大量实验表明,本框架性能优于现有纯视觉模型。在BBMP、铝箔和PCB数据集上,相较于YOLOv8/YOLOX,[email protected]:0.95分别提升+13.9、+1.4、+2.0个百分点,同时提升精确率与召回率。
原文摘要 · Abstract (English)
High inter-class similarity, extreme scale variation, and limited computational budgets hinder reliable visual recognition across diverse real-world data. Existing vision-centric and cross-modal approaches often rely on rigid fusion mechanisms and heavy annotation pipelines, leading to sub-optimal generalization. We propose the Distilled Large Language Model (LLM)-Driven Sparse Mixture-of-Experts (DS-MoE) framework, which integrates text-guided dynamic routing and lightweight multi-scale comprehension. The DS-MoE framework dynamically aligns textual semantics with defect-specific visual patterns through a sparse MoE architecture, where task-relevant experts are adaptively activated based on semantic relevance, resolving inter-class ambiguity. A lightweight MobileSAM encoder enables real-time inference while preserving multi-scale defect details. Extensive experiments on PCB, aluminum foil, and mold defect datasets demonstrate that our framework achieves superior performance compared to existing pure vision models. \textbf{DS-MoE} surpasses YOLOv8/YOLOX with gains of +13.9, +1.4, and +2.0 pp mAP@ 0.5:0.95 on BBMP, aluminum, and PCB, respectively, while also improving precision and recall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。