arXiv:2606.18389cs.CL2026-06

用激活值引导生成更高质量的低资源语言数据

Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation

论文配图:Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation
图 1 · 摘自论文原文
  • 通过调节模型早期层的激活值,控制生成文本的语言特征和质量
  • 在11种语言上提升数据多样性,低资源语言下游性能显著增强
  • 适合需要高效生成多样化低资源语言数据的研究者

大语言模型已成为合成数据生成的有效工具,尤其在低资源语言场景中,生成数据可提升下游任务表现。现有最优方法通常依赖目标语言示例的少样本提示,增加推理成本,并因词汇锚定降低数据多样性。本文探索激活值引导作为替代方案,研究两种策略:语言引导(聚焦语言身份)与质量引导(对比人工撰写与回译文本表示)。我们在四个开源LLM上、多个层级、11种语系多样的语言中评估,生成情感与主题分类数据并微调小型分类器。引导应用于零样本与少样本提示场景,对比非引导结果。结果显示,早期层引导能持续提升生成数据多样性,且常带来更强的下游性能,尤其在低资源语言中表现突出。

原文摘要 · Abstract (English)

Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated data can improve downstream task performance. Current best-performing approaches typically rely on few-shot prompting with target-language examples, which increases inference costs and may reduce diversity through lexical anchoring. In this work, we investigate activation steering as an alternative for low-resource synthetic data generation. We study two steering strategies: Language Steering, which targets the linguistic identity of a language, and Quality Steering, which captures well-formedness by contrasting human-written and backtranslated text representations. We evaluate these methods across four open-source LLMs, multiple layers, and 11 typologically diverse languages by generating sentiment and topic classification data and finetuning smaller classifiers. Steering is applied in both zero-shot and few-shot prompting settings and compared against non-steered counterparts. Our results show that steering on early layers consistently improves the diversity of generated data while often yielding stronger downstream model performance, particularly for low-resource languages.

合成数据低资源语言激活引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。