arXiv:2607.01433cs.AIcs.LG2026-07中稿 · ICML

不依赖数据或训练,通过权重调节提升大模型的发散思维能力

CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse

论文配图:CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
图 1 · 摘自论文原文
  • 通过对比权重调节实现无数据创造性增强
  • 在多项测试中提升原创性,最高达14个百分点
  • 可推广至开放任务,有效缓解模式崩溃问题

发散思维是创造力的关键,但大语言模型在开放问题上常生成相似回复,形成人工蜂群效应。本文提出CreativityNeuro,一种无需数据的对比权重调节方法,用于提升LLM的发散思维能力。在发散联想任务(DAT)中,该方法性能提升最高达14个人类百分点。在大规模人类评估(N=720)中,对替代用途测试(AUT)和任务任务的评估显示,CreativityNeuro显著提升原创性、惊喜度与创造力,且能迁移至长文本和更开放的任务。重要的是,三种任务中均显著降低模式崩溃程度。激活调节在DAT上表现相当,但无法迁移到AUT和任务任务,证明权重空间调节在泛化上的优势。结论:CreativityNeuro无需行为数据、重训练或基于梯度的微调,即可有效提升大模型在创意领域的表现。

原文摘要 · Abstract (English)

Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to open-ended questions, in what has been termed the artificial hivemind effect. Here, we introduce CreativityNeuro, a data-free method for enhancing divergent thinking in LLMs via contrastive weight steering. We evaluate our method across multiple creativity assessments and report several main findings. On the Divergent Association Task (DAT), a vocabulary-space creativity test, CreativityNeuro improves performance by up to 14 human percentile points. Next, in a large-scale human evaluation (N=720) on the Alternative Uses Test (AUT) and the Task Task, CreativityNeuro achieves significant improvements in originality, surprise, and creativity, transferring to longer-form and more open-ended tasks. Importantly, we find that across all three tasks, CreativityNeuro demonstrably reduces measures of mode collapse. Moreover, activation steering achieves comparable performance to CreativityNeuro on the DAT, but it does not transfer to the AUT and Task Task, demonstrating the effectiveness of weight-space steering in generalizing to unseen tasks. In conclusion, CreativityNeuro improves divergent thinking and reduces mode collapse without requiring behavioral data, re-training, or gradient-based fine-tuning, providing a straightforward way to enhance LLM performance in creative domains.

大模型创造力权重调节发散思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。