arXiv:2503.01622cs.CL2025-03ACL被引 22

构建大规模提示扰动数据集,评估大模型在多种提示变化下的表现稳定性。

DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation

  • 设计多维度提示扰动,生成上千种变体测试模型鲁棒性。
  • 发现少样本示例能降低模型对提示的敏感度,提升评估效率。
  • 适合关注模型评估可靠性与提示工程的研究者使用。

近期研究发现大语言模型对提示中的多种任意维度(如分隔符类型、答案编号方式、指令表述等)极为敏感,这质疑了单一提示评估方法的有效性。本文提出DOVE(Dataset Of Variation Evaluation),一个包含多个评估基准提示扰动的大规模数据集。与以往工作不同,我们从整体视角考察多种扰动的联合影响,每条样本生成数千种扰动变体。对多个模型家族进行评估后发现:存在高效选择高性能提示的方法;少样本示例可降低模型敏感性;部分实例在所有扰动下均表现困难。DOVE包含超过2.5亿个提示扰动及其模型输出,已公开发布,旨在推动社区开展更可靠、稳健且高效的模型评估。访问 https://slab-nlp.github.io/DOVE/ 浏览数据、贡献成果。

原文摘要 · Abstract (English)

Recent work found that LLMs are sensitive to a wide range of arbitrary prompt dimensions, including the type of delimiters, answer enumerators, instruction wording, and more. This throws into question popular single-prompt evaluation practices. We present DOVE (Dataset Of Variation Evaluation) a large-scale dataset containing prompt perturbations of various evaluation benchmarks. In contrast to previous work, we examine LLM sensitivity from an holistic perspective, and assess the joint effects of perturbations along various dimensions, resulting in thousands of perturbations per instance. We evaluate several model families against DOVE, leading to several findings, including efficient methods for choosing well-performing prompts, observing that few-shot examples reduce sensitivity, and identifying instances which are inherently hard across all perturbations. DOVE consists of more than 250M prompt perturbations and model outputs, which we make publicly available to spur a community-wide effort toward meaningful, robust, and efficient evaluation. Browse the data, contribute, and more: https://slab-nlp.github.io/DOVE/

大模型评估提示工程鲁棒性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。