用信息论方法评估模型对多领域评分任务的反应,揭示其偏好与不确定性。
Extending Minimal Pairs with Ordinal Surprisal Curves and Entropy Across Applied Domains
- 通过计算模型对评分选项的惊讶值,构建连续评分曲线。
- 在四个领域中,模型在预期位置出现明显最低惊讶值,熵值区分难易度。
- 无需生成文本,可直接分析模型对不同答案的置信程度。
最小配对范式通过对比模型对相反补全的预测概率,已被用于评估语言模型的语法知识,但主要局限于句法现象的二元合乎性判断。传统基于提示的评估需要昂贵的文本生成,可能诱发后验解释而非真实模型判断,并丢弃模型不确定性的信息。本文通过将基于惊讶值的评估从二元合乎性对比扩展到跨多个领域的序数分类与评分任务,解决了上述问题。不再要求模型生成答案,而是测量模型对评分量表(如1-5或1-9)上各位置的惊讶值(负对数概率),生成完整的惊讶值曲线,从而揭示模型的偏好及通过熵体现的不确定性。我们在四个领域进行了探索:社会-生态-技术系统分类、因果陈述识别(二元与分级)、隐喻语言检测,以及演绎定性编码。结果显示,惊讶值曲线在预期量表位置出现清晰极小值,且完成项上的熵能有效区分真正模糊项与较易项。
原文摘要 · Abstract (English)
The minimal pairs paradigm of comparing model probabilities for contrasting completions has proven useful for evaluating linguistic knowledge in language models, yet its application has largely been confined to binary grammaticality judgments over syntactic phenomena. Additionally, standard prompting-based evaluation requires expensive text generation, may elicit post-hoc rationalizations rather than model judgments, and discards information about model uncertainty. We address both limitations by extending surprisal-based evaluation from binary grammaticality contrasts to ordinal-scaled classification and scoring tasks across multiple domains. Rather than asking models to generate answers, we measure the information-theoretic "surprise" (negative log probability) they assign to each position on rating scales (e.g., 1-5 or 1-9), yielding full surprisal curves that reveal both the model's preferred response and its uncertainty via entropy. We explore this framework across four domains: social-ecological-technological systems classification, causal statement identification (binary and scaled), figurative language detection, and deductive qualitative coding. Across these domains, surprisal curves produce interpretable classification signals with clear minima near expected ordinal scale positions, and entropy over the completion tended to distinguish genuinely ambiguous items from easier items.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。