arXiv:2506.15629cs.CLcs.AI2025-06ACL被引 16

测试大模型按指令顺序生成句子的能力,发现其组合泛化仍不足。

Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

  • 设计新基准Ordered CommonGen,评估模型按序生成能力
  • 最先进模型仅75%有序覆盖率,说明指令遵循仍有缺陷
  • 适合研究大模型推理与指令理解的学者参考

在如CommonGen的生成式常识推理任务中,大语言模型需组合给定概念生成句子。但当提示包含概念顺序要求时,模型必须按指定顺序生成。为此,我们提出Ordered CommonGen基准,通过测量有序覆盖率来同时评估模型的组合泛化与指令遵循能力。对36个LLMs的全面分析显示,尽管模型总体理解指令意图,但仍存在对特定顺序模式的偏好,导致输出多样性低或在顺序变化时产生相同结果。即使是最符合指令的模型,有序覆盖率也仅约75%,表明在指令遵循与组合泛化方面仍需改进。

原文摘要 · Abstract (English)

In generative commonsense reasoning tasks such as CommonGen, generative large language models (LLMs) compose sentences that include all given concepts. However, when focusing on instruction-following capabilities, if a prompt specifies a concept order, LLMs must generate sentences that adhere to the specified order. To address this, we propose Ordered CommonGen, a benchmark designed to evaluate the compositional generalization and instruction-following abilities of LLMs. This benchmark measures ordered coverage to assess whether concepts are generated in the specified order, enabling a simultaneous evaluation of both abilities. We conducted a comprehensive analysis using 36 LLMs and found that, while LLMs generally understand the intent of instructions, biases toward specific concept order patterns often lead to low-diversity outputs or identical results even when the concept order is altered. Moreover, even the most instruction-compliant LLM achieved only about 75% ordered coverage, highlighting the need for improvements in both instruction-following and compositional generalization capabilities.

大模型指令遵循组合泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。