通过温度调整前的分布变化,能更准确预测大模型创意生成质量。
Before and After Temperature: A Distributional View of Creative LLM Generation
- 分析采样温度如何重塑令牌分布,捕捉创意生成的关键信号。
- 单个分词特征在创意排序上相关性达0.918,远超现有方法。
- 适合关注生成质量评估与分布动态机制的研究者。
大语言模型创意生成的无参考评估依赖困惑度、熵和最大概率差距等指标。我们发现更强的信号其实存在于采样前一步:温度对模型输出分布的重塑过程。在Llama-3.1-8B-Instruct对500个开放式创意提示(温度T ∈ {0.3, 0.8, 1.5})生成结果中,一个基于此重塑过程的单令牌特征,在与gpt-4o/gemini-2.5-pro平均评分(n=500)的相关性达到Spearman ρ=0.918,与三人人类多数评分(n=150)相关性为ρ=0.870。四个标准无参考基线(自困惑度、平均预测熵、最大概率差距、gzip压缩比)在两者上的最高相关性均约|ρ|=0.76,分别领先0.165和0.110,显著高于基线间差异。两组评分相关性ρ=0.83,超过人类评分上限ρ=0.77,说明评估不受评判噪声限制。机制上,这一优势源于非一致性区间的分布特征:当T=1.5时,累积质量宽度n₉₅(q)从~1增至~131,且后温度下质量脱离预温度前90%合理集约13个百分点。
原文摘要 · Abstract (English)
Reference-free evaluation of large language model (LLM) creativity relies on perplexity, entropy, and top-1 margin. We show that a much stronger signal lives one step earlier in the pipeline: in how sampling temperature \emph{reshapes} the model's token distribution before the next token is drawn. On Llama-3.1-8B-Instruct generations of 500 open-ended creative prompts at $T \in \{0.3, 0.8, 1.5\}$, a single per-token feature derived from this reshaping predicts the within-prompt creativity rank at Spearman $ρ{=}0.918$ against an averaged gpt-4o\,/\,gemini-2.5-pro judge ($n{=}500$) and $ρ{=}0.870$ against a three-rater human-majority ranking ($n{=}150$). Each of four standard reference-free baselines (self-perplexity, mean predictive entropy, top-1 margin, gzip compression ratio) tops out at $|ρ|\!\approx\!0.76$ on both ground truths: a gap of $+0.165$ on averaged-LLM and $+0.110$ on human-majority, both far larger than the spread among the baselines themselves. The two ground-truth panels agree with each other at $ρ{=}0.83$, above the inter-human ceiling of $ρ{=}0.77$, so the comparison is not bottlenecked by judge noise. Mechanistically, the win comes from a sharp distributional signature of the incoherence regime: at $T{=}1.5$ the cumulative-mass width $n_{95}(q)$ inflates from $\sim\!1$ to ${\sim}\!131$ tokens and post-temperature mass leaks off the pre-temperature top-$90\%$ plausible set by about $13$ percentage points. The per-token aggregates do not separate $T{=}0.8$ from $T{=}0.3$; discriminating the two coherent regimes is left to sequence-level features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。