用激活调控让模型更好说意大利语,效果不输微调
A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering
- 通过激活调控提升模型对意大利语的适配能力
- 意大利语生成质量与微调模型相当甚至更优
- 适合资源有限但需快速部署意大利语模型的场景
将模型适配到预训练数据中仅部分覆盖的语言,通常需要昂贵的微调。本文探索了基于激活调控的技术作为替代方案,以提升模型在意大利语任务上的表现。实验表明,意大利语激活调控(i)可适用于多种模型,(ii)在意大利语任务上性能达到或超过微调模型水平,(iii)生成内容的质量和一致性更高。同时讨论了在当前大模型已具备较高意大利语能力背景下,调控与微调各自的适用性。
原文摘要 · Abstract (English)
Adapting models to a language that was only partially present in the pre-training data requires fine-tuning, which is expensive in terms of both data and computational resources. As an alternative to fine-tuning, we explore the potential of activation steering-based techniques to enhance model performance on Italian tasks. Through our experiments we show that Italian steering (i) can be successfully applied to different models, (ii) achieves performances comparable to, or even better than, fine-tuned models for Italian, and (iii) yields higher quality and consistency in Italian generations. We also discuss the utility of steering and fine-tuning in the contemporary LLM landscape where models are anyway getting high Italian performances even if not explicitly trained in this language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。