剪枝专家模型在生物医学领域需权衡速度与事实可靠性
On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

- 针对生物医学任务设计领域特异性剪枝策略
- 适度剪枝可保持领域内性能,极端剪枝增加幻觉风险
- 高风险场景必须评估可靠性,不能只看效率提升
混合专家(MoE)模型通过选择性激活实现推理加速,但需常驻全部网络导致内存开销大。结构化专家剪枝是降低部署成本的实用方法。然而,以往研究多关注基准性能,对剪枝在高风险领域如生物医学中的事实可靠性影响研究不足。本文评估了四种MoE模型、六种剪枝方法及多种剪枝比例,在生成与分类任务中,涵盖领域内(生物医学)与跨领域设置。结果表明,适度剪枝可保持领域内任务性能且不立即引发可靠性下降,但极端剪枝会显著增加幻觉风险;当任务切换至通用领域时,性能与可靠性均迅速退化。这说明安全压缩高度依赖任务与领域特性。仅以性能评估剪枝后的MoE模型不足以支撑高风险应用,必须同时评估事实可靠性。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the whole network must remain loaded. Structured expert pruning is a practical approach for reducing deployment costs in resource-constrained settings. However, prior studies primarily evaluate benchmark utility, leaving the effect of pruning on factual reliability underexplored, particularly in high-stakes domains such as biomedicine. In this paper, we investigate how domain-specific expert pruning affects both utility and reliability. We assess four MoE models, six pruning methods, and multiple pruning ratios across generation and classification tasks under in-domain (biomedical) and cross-domain settings. Results reveal that moderate pruning preserves in-domain utility without immediate reliability decline, although hallucination risks increase at extreme pruning ratios. When shifting to the general domain, both utility and reliability degrade rapidly. These findings indicate that safe compression depends heavily on the task and domain. Evaluating pruned MoE models solely on utility is inadequate for high-stakes deployment without reliability assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。