用多维度创意信号优化大模型,生成更新颖多样内容。
Creative Preference Optimization
- 模块化注入多种创意维度,改进偏好优化目标。
- 在20万条人类生成数据上训练,超越GPT-4o表现。
- 适合追求高质量创意生成的开发者与研究者。
尽管大型语言模型在自然语言生成任务中表现优异,但在生成真正具有创新性、多样性、意外性和高质量内容方面仍显不足。现有方法多聚焦于单一维度或特定任务,难以通用化地提升创造力。本文提出创意偏好优化(CrPO),一种将多维度创意信号以模块化方式融入偏好优化目标的新方法。我们基于包含超过20万条人类生成回复及30项心理创造力评估的新型大规模人类偏好数据集MuCE,训练并评估了多个经过创意增强的模型。实验结果表明,所提模型在自动化和人类评估中均优于强基线,包括GPT-4o,生成内容更具新颖性、多样性和意外性,同时保持高输出质量。在NoveltyBench上的额外评测进一步验证了该方法的泛化能力。结果表明,在偏好框架内直接优化创造力是提升大模型创造力潜力的重要方向,且不损害输出质量。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated impressive performance across natural language generation tasks, their ability to generate truly creative content-characterized by novelty, diversity, surprise, and quality-remains limited. Existing methods for enhancing LLM creativity often focus narrowly on diversity or specific tasks, failing to address creativity's multifaceted nature in a generalizable way. In this work, we propose Creative Preference Optimization (CrPO), a novel alignment method that injects signals from multiple creativity dimensions into the preference optimization objective in a modular fashion. We train and evaluate creativity-augmented versions of several models using CrPO and MuCE, a new large-scale human preference dataset spanning over 200,000 human-generated responses and ratings from more than 30 psychological creativity assessments. Our models outperform strong baselines, including GPT-4o, on both automated and human evaluations, producing more novel, diverse, and surprising generations while maintaining high output quality. Additional evaluations on NoveltyBench further confirm the generalizability of our approach. Together, our results demonstrate that directly optimizing for creativity within preference frameworks is a promising direction for advancing the creative capabilities of LLMs without compromising output quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。