研究大模型如何把模糊程度词转成数字,发现效果受系统状态影响且差异变小。
Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

- 用10个程度副词测试模型在资源分配中的数值输出,控制变量隔离影响。
- 10个词压缩成5个输出值,强弱词区分明显,相关性高达0.845(p<0.001)。
- 系统接近极限时,弱词小调整、强词不动作、极端词冲顶,行为突变。
本文研究语言模型在需生成数值动作时是否保持程度词的顺序意义。基于Quirk等人程度修饰词分类体系,构建包含从"slightly"到"drastically"共10个英语程度副词的量表,在一个受控资源分配环境中,让Claude Haiku接收自然语言指令并输出数值分配,由确定性后端转化为可测量结果。唯一变量为程度词或初始系统状态,独立分析其对模型输出的影响。在T=0.0和T=0.7下共进行6,620次实验,发现三类模式:第一,模型将10个程度词压缩为5个不同中位数输出,低阶词均映射相同值,高阶词分属更高区间(斯皮尔曼等级相关rho = 0.845,p < 0.001);第二,当提供当前系统状态作为上下文时,按起始分配分组所捕获的秩次方差远高于按词分组(ε²基线=0.782 vs. ε²词=0.079),且随着系统逼近容量,词汇区分度降为零;第三,接近可行边界时,模型呈现三种行为模式:弱词小幅调整,强词完全回避,而"drastically"直接推向局部上限。这些模式在不同温度下持续存在,随机采样仅拓宽分布而不恢复词间序数差异。在该模型与场景下,模型对模糊程度词的数值解释呈压缩、状态依赖及近边界非连续特征。
原文摘要 · Abstract (English)
Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale of 10 English degree modifiers, from slightly to drastically, informed by the Quirk et al. degree-modifier taxonomy, in a controlled resource-allocation environment where Claude Haiku receives a natural-language instruction, produces a numeric allocation, and a deterministic backend converts that allocation into a measurable outcome. The only variable that changes between runs is the intensity word or the starting system state, isolating their effects on the model's numeric output. Across 6,620 runs at T=0.0 and T=0.7, three patterns emerge. First, the model compresses 10 intensity words into 5 distinct median outputs: four lower-tier words all map to the same value, while stronger words break into higher regimes (Spearman rho = 0.845, p < 0.001). Second, when the current system state is supplied as context, separate Kruskal-Wallis tests show that grouping by starting allocation captures far more rank-based variance than grouping by word (epsilon-squared baseline = 0.782 vs. epsilon-squared word = 0.079), and lexical differentiation collapses to zero as the system approaches capacity. Third, near feasibility limits the model exhibits three behavioral modes: weak words hedge with small adjustments, strong words abstain entirely, and the word drastically pushes to the local ceiling. These patterns persist across temperature, with stochastic sampling broadening distributions but not restoring ordinal distinctions between words. In this model and domain, the model's numeric interpretation of vague intensity words is compressed, state-dependent, and discontinuous near operational boundaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。