arXiv:2605.00195cs.LG2026-05被引 1

SFT让大模型生成变单一,新方法让回答更丰富且不失质量

Diversity in Large Language Models under Supervised Fine-Tuning

  • 提出TOFU损失函数,同时解决数据低频模式忽略和预训练知识遗忘问题
  • 实验证明SFT后生成多样性显著下降,TOFU可恢复并提升多样性30%以上
  • 适合追求生成多样性又不牺牲质量的模型优化场景

监督微调(SFT)对对齐大语言模型(LLMs)与用户意图至关重要,但普遍认为会抑制生成多样性。尽管这一现象常被提及,但缺乏系统的实证研究。已有方法从不同角度探讨了模型表达能力,但其视角差异提示需进一步深入。本研究将多样性下降归因于两个主要因素:微调数据集中低频模式被忽视,以及预训练知识的遗忘。基于理论分析,我们提出新型目标函数Tempered Focal (TOFU)损失,同时应对上述挑战。大规模评估证实,经SFT后生成广度明显缩小,而TOFU在多个模型与基准上显著提升输出多样性,同时保持高质量响应,为SFT提供了一种有原则的改进路径。

原文摘要 · Abstract (English)

Supervised Fine-Tuning (SFT) is essential for aligning Large Language Models (LLMs) with user intent, yet it is believed to suppress generative diversity. Although this reduction is frequently referenced, formal empirical testing of the phenomenon remains limited. The expressiveness of LLMs by itself was addressed by multiple prior methods. Their varying perspectives suggest that deeper investigation could yield further improvements. In this study, we attribute the decline to two primary drivers: the neglect of low-frequency patterns within fine-tuning datasets and the forgetting of preexisting knowledge. Motivated by our theoretical analysis, we develop Tempered Focal (TOFU) loss, a novel objective that addresses both stated challenges simultaneously. Our extensive evaluation confirms at scale that generation breadth narrows after SFT and strengthens the hypothesis explaining this effect. Across multiple models and benchmarks, we demonstrate that TOFU enhances output diversity while preserving high response quality, offering a principled approach to SFT.

大模型生成多样性微调TOFU

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。