arXiv:2503.22698cs.CL2025-03

提出GEM模型,在设备端平衡专精与通用性,提升多领域表现。

Fragile Mastery: Are Domain-Specific Trade-Offs Undermining On-Device Language Models?

  • 设计稀疏跨注意力路由机制,动态分配计算资源
  • 在四类设备上实现0.89跨域准确率,延迟低于100ms
  • 新评估指标揭示压缩越强模型越脆弱,适合边缘部署研究者

在资源受限的边缘设备上部署本地语言模型面临计算效率、内存、功耗与语言能力之间的多重权衡。本研究系统考察了领域特化优化与跨领域鲁棒性间的矛盾,提出通用边缘模型(GEM),通过稀疏跨注意力路由(SCAR)实现专精与泛化的协调。在八个领域(医疗、法律、金融、STEM、常识、对话AI、多语言、领域自适应)共47个基准测试中,传统优化方法使目标任务困惑度降低18%-25%,但导致通用任务性能骤降,F1下降12%-29%。GEM在树莓派4、像素6、iPhone 13及定制神经处理单元(NPUs)上实现跨域F1达0.89,延迟低于100ms。相比GPT-4 Lite,GEM在通用任务上提升7%,且保持领域特定性能相当。提出三个新度量工具:领域特化指数(DSI)、泛化差距(GG)、跨域迁移比(CDTR),显示模型压缩强度与脆弱性呈强相关。

原文摘要 · Abstract (English)

The application of on-device language models (ODLMs) on resource-constrained edge devices is a multi-dimensional problem that strikes a fine balance between computational effectiveness, memory, power usage, and linguistic capacity across heterogeneous tasks. This holistic study conducts a thorough investigation of the trade-offs between domain-specific optimization and cross-domain robustness, culminating in the proposal of the Generalized Edge Model (GEM), a new architecture that aims to balance specialization and generalization in a harmonious manner. With a rigorous experimental approach testing 47 well-chosen benchmarks in eight domains--healthcare, law, finance, STEM, commonsense, conversational AI, multilingual, and domain-adaptive tasks--we show that conventional optimization techniques decrease target task perplexity by 18-25% but result in a precipitous decline in general-task performance with F1 scores decreasing by 12-29%, as reported by Liu et al. GEM employs a Sparse Cross-Attention Router (SCAR) to dynamically allocate computation to a variable number of computing resources with a cross-domain F1 accuracy of 0.89 on less than 100ms latency across Raspberry Pi 4, Pixel 6, iPhone 13, and bespoke custom neural processing units (NPUs). Compared to GPT-4 Lite, GEM enhances the general-task level by 7% with respect and parity in domain-specific performance. We propose three new measurement tools--Domain Specialization Index (DSI), Generalization Gap (GG), and Cross-Domain Transfer Ratio (CDTR)--which show strong correlation between model compression intensity and brittleness.

边缘计算语言模型模型压缩多领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。