arXiv:2503.17126cs.CLcs.LG2025-03被引 44

让大模型生成更多样化的创意写作,同时保持高质量。

Modifying Large Language Model Post-Training for Diverse Creative Writing

  • 在训练目标中引入样本偏离度,鼓励学习稀有优质输出。
  • 80亿参数模型在多样性上媲美人工数据集,质量接近GPT-4o。
  • 适合需要多样化创意内容的场景,如文案、故事生成。

由于创意写作任务无唯一正确答案,训练大语言模型时应能生成多样且有效的输出。然而,现有后训练方法多关注生成质量,忽视了输出多样性。为此,本文研究如何通过后训练提升创意写作的多样性与质量。核心思路是将“偏离度”——同一提示下样本与其他样本的差异程度——纳入训练目标,以促进对稀有高质量样本的学习。基于直接偏好优化(DPO)和几率比偏好优化(ORPO),实验表明该方法可在不显著降低质量的前提下提升输出多样性。最佳80亿参数模型在多样性上达到人类创作数据集水平,质量接近我们评估的最优指令微调模型GPT-4o与DeepSeek-R1。通过人工评估、消融实验及与已有方法DivPO的对比,进一步验证了本方法的有效性。

原文摘要 · Abstract (English)

As creative writing tasks do not have singular correct answers, large language models (LLMs) trained to perform these tasks should be able to generate diverse valid outputs. However, LLM post-training often focuses on improving generation quality but neglects to facilitate output diversity. Hence, in creative writing generation, we investigate post-training approaches to promote both output diversity and quality. Our core idea is to include deviation -- the degree of difference between a training sample and all other samples with the same prompt -- in the training objective to facilitate learning from rare high-quality instances. By adopting our approach to direct preference optimization (DPO) and odds ratio preference optimization (ORPO), we demonstrate that we can promote the output diversity of trained models while minimally decreasing quality. Our best model with 8B parameters could achieve on-par diversity as a human-created dataset while having output quality similar to the best instruction-tuned models we examined, GPT-4o and DeepSeek-R1. We further validate our approaches with a human evaluation, an ablation, and a comparison to an existing diversification approach, DivPO.

创意写作多样性生成后训练大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。