arXiv:2504.07965cs.LGcs.CL2025-04被引 2

测试32个语言模型对词语相似度的判断,发现小模型也能接近人类水平。

Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments

  • 用词三元组任务评估模型语义表示与人类判断的匹配度
  • 指令微调后模型与人类判断一致性显著提升
  • 大模型在行为和表征层面都更接近人类,小模型则不一致

小型和中型生成式语言模型受到越来越多关注。其规模和可获取性使其可在行为与表征层面进行分析,从而探究两者的交互机制。我们评估了32个公开可用的语言模型在词三元组任务中与人类相似性判断的表征和行为对齐情况。这一新评测设置超越了常见的成对比较,用于探查语言中的语义关联。研究发现:(1)即使小型语言模型的表征也能达到人类水平的对齐;(2)指令微调的模型变体表现出显著更高的一致性;(3)各层对齐模式高度依赖于具体模型;(4)基于模型行为响应的对齐程度强烈依赖于模型规模,仅在最大模型上才与表征对齐相匹配。

原文摘要 · Abstract (English)

Small and mid-sized generative language models have gained increasing attention. Their size and availability make them amenable to being analyzed at a behavioral as well as a representational level, allowing investigations of how these levels interact. We evaluate 32 publicly available language models for their representational and behavioral alignment with human similarity judgments on a word triplet task. This provides a novel evaluation setting to probe semantic associations in language beyond common pairwise comparisons. We find that (1) even the representations of small language models can achieve human-level alignment, (2) instruction-tuned model variants can exhibit substantially increased agreement, (3) the pattern of alignment across layers is highly model dependent, and (4) alignment based on models' behavioral responses is highly dependent on model size, matching their representational alignment only for the largest evaluated models.

语言模型语义对齐行为评估模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。