提出Spectrum Tuning方法,提升模型在多样化输出任务中的可控性和覆盖面。
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
- 基于Spectrum Suite数据集进行后训练,增强模型上下文可控性
- 相比传统微调,显著提升输出空间覆盖率与分布对齐度
- 适合需要多样输出的场景,如创意写作、偏好控制等
语言模型后训练虽提升了指令遵循和下游任务表现,却常忽视对多有效答案任务的负面影响。在创意写作、合成数据生成或偏好引导等任务中,模型需覆盖完整输出分布而非单一正确答案。本文提出三个理想特性:上下文可调控性、有效输出空间覆盖度、分布对齐性,并通过三类模型验证现有后训练方法会削弱这些能力。区分了两种上下文学习:一种是激发已有知识,另一种是利用上下文信息覆盖先验、实现新分布引导。为此构建Spectrum Suite,涵盖40+数据源、90+任务,覆盖人类偏好、数值分布等多种分布。实验发现,当前方法虽能激发潜在能力,但损害上下文可调控性。为此提出Spectrum Tuning,使用该数据集优化后训练,在保持性能的同时显著提升可调控性、输出空间覆盖度及分布对齐性。
原文摘要 · Abstract (English)
Language model post-training has enhanced instruction-following and performance on many downstream tasks, but also comes with an often-overlooked cost on tasks with many possible valid answers. On many tasks such as creative writing, synthetic data generation, or steering to diverse preferences, models must cover an entire distribution of outputs, rather than a single correct answer. We characterize three desiderata for conditional distributional modeling: in-context steerability, valid output space coverage, and distributional alignment, and document across three model families how current post-training can reduce these properties. In particular, we disambiguate between two kinds of in-context learning: ICL for eliciting existing underlying knowledge or capabilities, and in-context steerability, where a model must use in-context information to override its priors and steer to a novel data generating distribution. To better evaluate and improve these desiderata, we introduce Spectrum Suite, a large-scale resource compiled from >40 data sources and spanning >90 tasks requiring models to steer to and match diverse distributions ranging from varied human preferences to numerical distributions and more. We find that while current post-training techniques elicit underlying capabilities and knowledge, they hurt models' ability to flexibly steer in-context. To mitigate these issues, we propose Spectrum Tuning, a post-training method using Spectrum Suite to improve steerability and distributional coverage. We find that Spectrum Tuning often improves over pretrained and typical instruction-tuned models, enhancing steerability, spanning more of the output space, and improving distributional alignment on held-out datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。