用几何方法揭示大模型在创作中的不确定性差距
The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models
- 用可信集分析文本生成的不确定性分布
- 最佳模型仅达0.434的人类创意匹配度
- 解码策略影响近四成至七成的信念不确定性
理解大语言模型中的不确定性仍是核心挑战,尤其在存在多种合理输出的创造性任务中。本文提出一种基于可信集(即概率分布凸包)的几何框架,量化并分解神经文本生成中的不确定性,并以人类创意差异为校准基准。在WritingPrompts数据集上,针对500个创作提示(每个有10个独立人类续写),评估四种语言模型在五种解码策略下的表现,共生成10万条故事。可信集分析显示,模型与人类创意间的差距显著,最优模型(Gemma-2B,温度0.7)的校准得分仅为0.434。我们将总不确定性分解为认知型和随机型两部分,发现解码策略贡献了39.4%至72.0%的认知不确定性。模型规模与校准质量相关性弱,基础模型与指令微调模型间无显著差异。该几何框架为提升人机共创一致性提供了可操作洞察。实验框架已全部开源。
原文摘要 · Abstract (English)
Understanding uncertainty in large language models remains a fundamental challenge, particularly in creative tasks where multiple valid outputs exist. We present a geometric framework using credal sets - convex hulls of probability distributions - to quantify and decompose uncertainty in neural text generation, calibrated against human creative variation. Analyzing 500 creative writing prompts from the WritingPrompts dataset with 10 unique human continuations each, we evaluate four language models across five decoding strategies, generating 100,000 stories. Our credal set analysis reveals substantial gaps in capturing human creative variation, with the best model-human calibration reaching only 0.434 (Gemma-2B with temperature 0.7). We decompose total uncertainty into epistemic and aleatoric components, finding that the choice of decoding strategy contributes 39.4% to 72.0% of total epistemic uncertainty. Model scale shows weak correlation with calibration quality and no significant difference exists between base and instruction-tuned models in calibration quality. Our geometric framework provides actionable insights for improving generation systems for human-AI creative alignment. We release our complete experimental framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。