arXiv:2609.02760cs.AI2026-09

通过智能选子网,让工厂本地助手在低功耗设备上高效运行。

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

  • 基于检索增强和结构压缩,按质量与性能选最优子网部署。
  • 子网在1.3~5瓦功耗下运行,恢复原模型85.4%的问答质量。
  • 适合资源受限的工业现场部署,兼顾速度、功耗与准确率。

本地化助手可为工厂工人提供机器文档的对话访问,但具备该能力的模型通常无法适配产线硬件。我们发现,在结构压缩与检索引导的适应后,模型规模不再可靠预测适配后的回答质量:通用能力几乎随参数量线性下降,而评估的检索增强型回答质量却不呈此趋势。因此,将部署视为适应后的子网络选择问题,针对每台设备根据评估的回答质量与实测设备吞吐量,在可配置的通用能力下限与内存预算约束下选择子网络;仅优化大小、速度或质量之一都会牺牲能力或吞吐量。采用夹心式就地蒸馏训练的权重共享超网络,使该选择过程成本低廉。在制造手册案例研究中,模型提取导致13.7%的质量损失,而检索增强蒸馏使其恢复至原始质量的95.4%,挽回了约三分之二的损失,同一助手可在三类异构边缘设备上以1.3至5瓦待机功耗运行。

原文摘要 · Abstract (English)

On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable predictor of adapted answer quality: general capability falls almost linearly with parameter count, while judged retrieval-augmented answer quality does not. We therefore treat deployment as a post-adaptation selection problem, committing one sub-network per device on judged answer quality and measured on-device throughput under a configurable general-capability floor and memory budget; rules that optimize size, speed, or quality alone each give up capability or throughput. A weight-shared supernetwork trained with sandwich-style in-place distillation keeps this selection inexpensive. In a manufacturing-manual case study, extraction costs 13.7 percent of the unpruned model's judged quality and retrieval-grounded distillation returns it to within 4.6 percent, recovering two thirds of the loss, and the same assistant runs across three heterogeneous edge tiers at 1.3 to 5 watts standby.

边缘计算检索增强模型压缩工业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。