提出词汇覆盖度评分,量化采样策略如何压制低频高信息词
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)

- 设计词汇覆盖度评分(WCS),衡量采样过滤器对人类词汇的数学剔除程度
- 发现行业标准采样参数会大幅降低低频词存活率,导致语言同质化
- 适合关注生成文本多样性、模型可解释性的研究者和开发者
现代大语言模型常因生成重复、单调的文本而受到批评,尽管其潜在词汇量巨大。现有研究多聚焦于模型知识与训练数据,本文则考察解码机制对语言多样性的抑制作用。我们提出词汇覆盖度评分(WCS),用于量化标准采样滤波器(如Top-$p$、Top-$k$、Min-$p$)对上下文恰当的人类词汇的数学剔除程度。WCS不评估静态知识,而是衡量低频、高信息量词汇在采样参数下的词汇存活率。通过对人类撰写语料片段中的开源模型进行审计,我们识别出哪些语义上合理的词汇选择因解码器而不可达,即使它们存在于概率空间中。结果表明,行业标准采样默认设置实质上成为非故意的审查机制,将人类表达的独特纹理平滑为同质化话语。WCS为优化文本连贯性与词汇丰富性的权衡提供了严谨框架,是保留生成模型中人类语言多样性的诊断工具。
原文摘要 · Abstract (English)
Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. While previous research has focused on model knowledge and training data, we investigate the role of decoding mechanics in suppressing linguistic diversity. We introduce the Word Coverage Score (WCS), a metric that quantifies the extent to which contextually appropriate human vocabulary is mathematically pruned by standard sampling filters (e.g., Top-$p$, Top-$k$, and Min-$p$). Rather than assessing static knowledge, the WCS measures the lexical survival rate of low-frequency, high-information human words as a function of sampling parameters. By auditing open-weight models on human-authored corpus fragments, we identify which logical lexical choices are rendered unreachable by the decoder, even when they reside within the probability space. Our results provide quantitative evidence that industry-standard sampling defaults act as unintended censorship mechanisms, smoothing the unique textures of human expression into a homogenized discourse. The WCS offers a rigorous framework for optimizing the trade-off between text coherence and lexical richness, providing a diagnostic tool for preserving the diversity of human language in generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。