arXiv:2508.02913cs.AI2025-08

用推理向量提升日语大模型能力,无需大量资源。

Enhancing Japanese Large Language Models with Reasoning Vectors

  • 从推理模型提取权重变化向量,注入日语模型
  • 仅需少量资源即实现性能显著提升
  • 方法简单高效,适合低资源语言模型优化

后训练方法已显著提升主流大语言模型的性能与推理能力,但日语大模型因资源限制难以实现类似效果。受任务向量启发——通过训练前后权重变化捕捉特定任务的改进——我们从推理型大模型中提取推理向量,并将其应用于日语大模型以增强其表现。尽管资源有限,我们提出一种简单有效的方案,实现了显著性能提升,有望为其他低资源语言提供借鉴。

原文摘要 · Abstract (English)

Post-training methods have improved the performance and enhanced the reasoning capability for mainstream large language models (LLMs), but the same is challenging for Japanese LLMs to achieve due to the amount of resources required. Inspired by task vectors that extract the change of weights before and after training, specifically for a certain task, we obtain reasoning vectors from reasoning LLMs and apply them to Japanese LLMs to boost their performance. While the resources available present a challenge to improve Japanese LLMs, we present a simple and effective way to obtain high improvement and hope to inspire for other languages.

日语模型推理向量低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。