揭示大模型创作输出中提示词、模型和采样噪声的贡献比例。
Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks
- 12个模型在10个创意任务上各生成100样本,量化三类因素影响
- 提示词对创意质量影响达36.43%,接近模型差异的40.94%
- 单次采样易被噪声干扰,需多次采样才可信
LLM输出变异性的来源是什么?我们通过在10个创意任务上评估12个大模型,每个生成100个样本(共12,000条),分析提示词、模型选择与采样随机性的影响。对于输出质量(原创性),提示词解释了36.43%的方差,与模型选择(40.94%)相当;但对于输出数量(流畅性),模型选择(51.25%)和模型内方差(33.70%)占主导,提示词仅解释4.22%。提示词对质量有显著调控作用,但模型内采样波动(10%-34%)不容忽视,单次采样评估可能将随机噪声误判为真实效应。
原文摘要 · Abstract (English)
How much of LLM output variance is explained by prompts versus model choice versus stochasticity through sampling? We answer this by evaluating 12 LLMs on 10 creativity prompts with 100 samples each (N = 12,000). For output quality (originality), prompts explain 36.43% of variance, comparable to model choice (40.94%). But for output quantity (fluency), model choice (51.25%) and within-LLM variance (33.70%) dominate, with prompts explaining only 4.22%. Prompts are powerful levers for steering output quality, but given the substantial within-LLM variance (10-34%), single-sample evaluations risk conflating sampling noise with genuine prompt or model effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。