通过删减无关语言词汇,让大模型更专注韩语任务,提升稳定性与性能。
Optimizing Korean-Centric LLMs via Token Pruning
- 针对韩语任务,删减非必要语言的词元和嵌入参数,实现模型压缩。
- 在机器翻译等韩语任务上,性能显著提升,生成更稳定。
- 适合资源受限场景下的韩语专用模型部署,尤其关注推理效率。
本文系统评估了基于词元剪枝(token pruning)优化的先进多语言大模型在韩语为中心自然语言处理任务中的表现。该压缩技术通过移除与目标应用无关的语言对应词元和嵌入参数,降低模型复杂度。实验聚焦 Qwen3、Gemma-3、Llama-3、Aya 等架构,在原始、英韩(EnKo)、英韩中(EnKoZh)三种词汇配置下,使用通用能力、文化素养、指令遵循及机器翻译等基准进行评测。结果表明,词元剪枝能有效消除语言混淆,显著提升生成稳定性;在韩语任务中,机器翻译性能普遍增强。尽管指令遵循能力受模型架构影响存在差异,但词汇量大幅缩减验证了该方法在内存受限、领域专用部署中的高效性,仅带来轻微推理延迟降低。
原文摘要 · Abstract (English)
This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression technique that eliminates tokens and embedding parameters corresponding to languages irrelevant to the target application. Focusing on Korean-centric natural language processing (NLP) tasks, we evaluate architectures including Qwen3, Gemma-3, Llama-3, and Aya across three vocabulary configurations: Original, English-Korean (EnKo), and English-Korean-Chinese (EnKoZh). Performance is assessed using established benchmarks for general aptitude, cultural literacy, instruction following, and machine translation. Our findings indicate that token pruning significantly improves generation stability by eliminating language confusion, and in the case of machine translation, frequently enhances performance on Korean-specific tasks. While instruction-following capabilities display architecture-dependent variance linked to latent cross-lingual representations, the significant reduction in vocabulary size validates token pruning as a highly effective optimization strategy for memory-constrained, domain-specific deployments, despite modest gains in inference latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。