arXiv:2506.13681cs.CLcs.LG2025-06被引 7

重新评估 min-p 采样,发现其并未优于传统方法。

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models

  • 复现原论文实验,发现其数据与统计方法存在错误
  • 在多个基准测试中,min-p 未超越基础采样方法
  • 揭露其社区采用数据造假,结论不可信

语言模型的采样策略影响生成内容的质量与多样性,对研究与应用均有重要影响。Nguyen 等人(2024)提出的 min-p 采样声称在创造性与连贯性上优于 basic、top-k、top-p 等经典采样方法,其工作被 ICLR 2025 评为第18高分提交,并入选口头报告。本文对原论文的四条核心证据进行系统重检:首先,原论文的人工评估缺失数据、统计方法错误且反馈描述失实;我们重分析表明 min-p 在质量、多样性及二者权衡上均未优于基线;原作者虽用新任务与评分标准重测,仍未能证明优势。其次,全面扫描原论文的 NLP 基准测试发现,在控制超参数数量的前提下,min-p 未超越基线。第三,原论文的 LLM-as-a-Judge 评估方法不清晰,报告不一致。第四,原论文宣称的 49,000 个 GitHub 仓库与 110 万星标被证实无据,已被移除;即使修正后仍具误导性。结论:原论文所提证据不足以支持 min-p 能提升质量、多样性或二者的平衡。

原文摘要 · Abstract (English)

Sampling from language models impacts the quality and diversity of outputs, affecting both research and real-world applications. Recently, Nguyen et al. 2024's "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs" introduced a new sampler called min-p, claiming it achieves superior quality and diversity over established samplers such as basic, top-k, and top-p sampling. The significance of these claims was underscored by the paper's recognition as the 18th highest-scoring submission to ICLR 2025 and selection for an Oral presentation. This paper conducts a comprehensive re-examination of the evidence supporting min-p and reaches different conclusions from the original paper's four lines of evidence. First, the original paper's human evaluations omitted data, conducted statistical tests incorrectly, and described qualitative feedback inaccurately; our reanalysis demonstrates min-p did not outperform baselines in quality, diversity, or a trade-off between quality and diversity; in response to our findings, the authors of the original paper conducted a new human evaluation using a different implementation, task, and rubric that nevertheless provides further evidence min-p does not improve over baselines. Second, comprehensively sweeping the original paper's NLP benchmarks reveals min-p does not surpass baselines when controlling for the number of hyperparameters. Third, the original paper's LLM-as-a-Judge evaluations lack methodological clarity and appear inconsistently reported. Fourth, community adoption claims (49k GitHub repositories, 1.1M GitHub stars) were found to be unsubstantiated, leading to their removal; the revised adoption claim remains misleading. We conclude that evidence presented in the original paper fails to support claims that min-p improves quality, diversity, or a trade-off between quality and diversity.

采样方法大模型可信评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。