arXiv:2503.06573cs.CLcs.AI2025-03中稿 · the 5th Workshop o…被引 11

构建7000条真实用户多约束指令数据集,评估大模型在复杂场景下的指令遵循能力。

WildIFEval: Instruction Following in the Wild

  • 收集7000条真实用户指令,涵盖8类高阶约束类型。
  • 大模型在多约束任务上表现不佳,仍有巨大提升空间。
  • 揭示约束数量与类型对模型表现的影响规律,适合评测与优化研究。

近期大语言模型在遵循用户指令方面表现出色,但处理具有多重约束的指令仍是重大挑战。本文提出WildIFEval——一个包含7000条真实用户指令的大规模数据集,覆盖广泛的语言和主题范围,且包含多样化的多约束条件。不同于以往数据集,本数据集的约束来自自然用户输入,并被归纳为8个高层类别,以反映现实场景中的分布与动态。基于WildIFEval,我们对主流大模型进行了全面评测,结果表明该数据集能有效区分小模型与大模型的表现,且所有模型在该任务上均有显著提升空间。我们进一步分析了约束数量与类型对性能的影响,揭示了模型在多约束情境下的行为模式。数据集已公开,旨在推动复杂真实条件下指令遵循的研究。

原文摘要 · Abstract (English)

Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we introduce WildIFEval - a large-scale dataset of 7K real user instructions with diverse, multi-constraint conditions. Unlike prior datasets, our collection spans a broad lexical and topical spectrum of constraints, extracted from natural user instructions. We categorize these constraints into eight high-level classes to capture their distribution and dynamics in real-world scenarios. Leveraging WildIFEval, we conduct extensive experiments to benchmark the instruction-following capabilities of leading LLMs. WildIFEval clearly differentiates between small and large models, and demonstrates that all models have a large room for improvement on such tasks. We analyze the effects of the number and type of constraints on performance, revealing interesting patterns of model constraint-following behavior. We release our dataset to promote further research on instruction-following under complex, realistic conditions.

指令遵循大模型评测多约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。