arXiv:2505.00047cs.CL2025-05被引 51

基础模型在随机与创意任务上优于对齐模型。

Base Models Beat Aligned Models at Randomness and Creativity

  • 对比基础模型与对齐模型在随机和创意任务中的表现
  • 对齐模型在生成随机数时偏好特定数字,游戏策略易被预测
  • 适合需要不可预测性或原创性的场景,如艺术创作、博弈设计

对齐已成为大语言模型开发的默认要素,例如基于人类反馈的强化学习使模型更安全、更遵从指令并在复杂任务中表现更好。然而我们提出,这些技术不应普遍应用,并展示了多种任务中基础语言模型始终优于主流对齐版本。特别关注需要不可预测输出的任务,如随机数生成、混合策略游戏(石头剪刀布、捉迷藏)以及创意写作。在每项任务中,对齐模型趋向于狭窄行为,导致明显劣势:例如倾向于生成“7”而非均匀分布的随机数,在某些游戏状态下几乎完全可预测,或优先选择悦耳表达而非创造性原创。在测试的多个模型中,基准测试表现越好,反而在我们的任务中表现越差,表明所需能力之间存在有效权衡。

原文摘要 · Abstract (English)

Alignment has quickly become a default ingredient in LLM development, with techniques such as reinforcement learning from human feedback making models act safely, follow instructions, and perform ever-better on complex tasks. While these techniques are certainly useful, we propose that they should not be universally applied and demonstrate a range of tasks on which base language models consistently outperform their popular aligned forms. Particularly, we study tasks that require unpredictable outputs, such as random number generation, mixed strategy games (rock-paper-scissors and hide-and-seek), and creative writing. In each case, aligned models tend towards narrow behaviors that result in distinct disadvantages, for instance, preferring to generate "7" over other uniformly random numbers, becoming almost fully predictable in some game states, or prioritizing pleasant writing over creative originality. Across models tested, better performance on common benchmarks tends to correlate with worse performance on our tasks, suggesting an effective trade-off in the required capabilities.

大模型对齐随机性创意生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。