arXiv:2510.12699cs.CLcs.AI2025-10被引 2

提出生成空间大小概念,解决大模型创作同质与事实幻觉问题

Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations

  • 用生成空间大小衡量模型对提示的语义输出多样性
  • EigenScore等内部指标能准确识别幻觉,优于传统多样性评估
  • 可检测提示模糊、解释推理偏差、引导模型生成更丰富结果

不同开放式生成任务需要不同程度的输出多样性。然而当前大模型常出现偏差:在创意任务中输出过于单一,在事实类任务中则产生多样但错误的回答。本文认为这两种问题本质统一于有效生成空间大小(GSS)——即模型针对提示考虑的语义上不同的输出集合。我们构建了GSSBench,一套包含具有真实GSS关系的提示对的任务集,用于评估各类度量方法并分析模型行为偏差。实验发现,仅使用模型内部信息的幻觉检测指标(尤其是EigenScore)显著优于标准多样性与不确定性度量,且提供可解释的内部表征洞察。本文展示了GSS的三项应用:(1) 检测提示模糊性并预测澄清问题以增强对齐;(2) 解释推理模型中的过度思考与思考不足现象;(3) 驱动模型扩展生成空间,实现高质量且多样的输出。

原文摘要 · Abstract (English)

Different open-ended generation tasks require different degrees of output diversity. However, current LLMs are often miscalibrated. They collapse to overly homogeneous outputs for creative tasks and hallucinate diverse but incorrect responses for factual tasks. We argue that these two failure modes are unified by, and can both be addressed by, the notion of effective generation space size (GSS) -- the set of semantically distinct outputs a model considers for a prompt. We present GSSBench, a task suite of prompt pairs with ground-truth GSS relationships to assess different metrics and understand where models diverge from desired behavior. We find that hallucination detection metrics, particularly EigenScore, consistently outperform standard diversity and uncertainty quantification metrics, while using only model internals, providing interpretable insights into a model's internal task representations. We demonstrate three applications of GSS: (1) detecting prompt ambiguity and predicting clarification questions for better grounding, (2) interpreting overthinking and underthinking in reasoning models, and (3) steering models to expand their generation space to yield high-quality and diverse outputs.

大模型生成生成空间幻觉检测提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。