为大模型社会认知评估建立理论框架,避免误判能力广度。
Theory Trace Card: Theory-Driven Socio-Cognitive Evaluation of LLMs
- 提出理论溯源卡,明确定义评估目标与能力构成
- 揭示当前评估中因缺乏理论支撑导致的过度推断问题
- 适合评估设计者与审稿人用于提升评测可信度
大语言模型的社会认知评测常无法预测真实表现,即使模型在评测中得分很高。现有研究将此差距归因于测量与效度问题,但我们认为更根本的问题在于:多数评测未明确目标能力的理论定义,使任务表现与真实能力之间的关联变得隐含。这导致仅覆盖能力局部的任务被误认为证明了全面能力,形成系统性效度幻觉。为此,本文提出两个贡献:首先,诊断并形式化该理论缺失为根本性缺陷,引发对评测结果的系统性误读;其次,提出理论溯源卡(Theory Trace Card, TTC),一种轻量级文档工具,用于显式说明评测的理论基础、能力组成、操作化方式及局限性。TTC通过完整呈现从理论到评分的验证链条,提升评测的可解释性与可复用性,无需修改基准或统一理论标准。
原文摘要 · Abstract (English)
Socio-cognitive benchmarks for large language models (LLMs) often fail to predict real-world behavior, even when models achieve high benchmark scores. Prior work has attributed this evaluation-deployment gap to problems of measurement and validity. While these critiques are insightful, we argue that they overlook a more fundamental issue: many socio-cognitive evaluations proceed without an explicit theoretical specification of the target capability, leaving the assumptions linking task performance to competence implicit. Without this theoretical grounding, benchmarks that exercise only narrow subsets of a capability are routinely misinterpreted as evidence of broad competence: a gap that creates a systemic validity illusion by masking the failure to evaluate the capability's other essential dimensions. To address this gap, we make two contributions. First, we diagnose and formalize this theory gap as a foundational failure that undermines measurement and enables systematic overgeneralization of benchmark results. Second, we introduce the Theory Trace Card (TTC), a lightweight documentation artifact designed to accompany socio-cognitive evaluations, which explicitly outlines the theoretical basis of an evaluation, the components of the target capability it exercises, its operationalization, and its limitations. We argue that TTCs enhance the interpretability and reuse of socio-cognitive evaluations by making explicit the full validity chain, which links theory, task operationalization, scoring, and limitations, without modifying benchmarks or requiring agreement on a single theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。