用'模因'视角重新评估大模型,揭示传统方法忽略的能力结构。
Probing Memes in LLMs: A Paradigm for the Entangled Evaluation World
- 将大模型视为由可复制的'模因'构成,通过感知矩阵捕捉模型与数据交互
- 在4507个模型上发现精英模型反而在简单问题上失败的反直觉现象
- 适合关注模型群体行为差异的研究者和评测系统设计者
当前大语言模型(LLMs)的评估范式将模型与数据分开处理,仅提供粗略描述:数据集中的条目被视为预标注项,模型则以整体得分(如准确率)概括,忽略了模型在不同属性数据上的行为多样性。为填补这一空白,本文将LLM视为由模因构成的系统——借鉴道金斯提出的文化基因概念,用于复制知识与行为。基于此,提出Probing Memes范式,将评估重构为模型与数据交织的复杂世界。该范式核心是感知矩阵,可捕捉模型-条目交互,进而支持探测属性(用于刻画条目特征)和模因分数(用于描绘模型行为特征)。在9个数据集和4,507个模型上的应用表明,该方法揭示了传统范式下不可见的能力结构,并量化了隐藏现象(如顶尖模型在多数模型能答对的问题上失败)。该方法不仅支持更丰富、可扩展的基准测试,也实现了基于群体的大模型评估。
原文摘要 · Abstract (English)
Current evaluation paradigms for large language models (LLMs) characterize models and datasets separately, yielding coarse descriptions: items in datasets are treated as pre-labeled entries, and models are summarized by overall scores such as accuracy, together ignoring the diversity of population-level model behaviors across items with varying properties. To address this gap, this paper conceptualizes LLMs as composed of memes, a notion introduced by Dawkins as cultural genes that replicate knowledge and behavior. Building on this perspective, the Probing Memes paradigm reconceptualizes evaluation as an entangled world of models and data. It centers on a Perception Matrix that captures model-item interactions, enabling Probe Properties for characterizing items and Meme Scores for depicting model behavioral traits. Applied to 9 datasets and 4,507 LLMs, Probing Memes reveals hidden capability structures and quantifies phenomena invisible under traditional paradigms (e.g., elite models failing on problems that most models answer easily). It not only supports more informative and extensible benchmarks but also enables population-based evaluation of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。