用理论框架验证大模型能力测评是否靠谱
Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
- 引入诺莫逻辑网络,系统检验大模型能力评估的合理性
- 指出当前评测方法缺乏对能力本质的理论支撑
- 适合关注模型评估可信度的研究者和评测设计者
近年来机器学习领域越来越多地基于基准测试表现,将推理、心智理论等类人能力归因于大语言模型(LLMs)。本文从建构效度视角审视这一做法,即如何将理论上的能力与实证测量相联系。文章对比了三种重要框架:Cronbach和Meehl提出的诺莫逻辑框架、Messick提出并由Kane完善的推断框架,以及Borsboom的因果框架。认为诺莫逻辑框架最适合作为当前大模型能力研究的基础,它避免了因果框架的强本体论承诺,同时比推断框架提供了更丰富的建构意义阐述。通过推理能力评估的案例,探讨采纳诺莫逻辑框架对大模型研究的深层含义。
原文摘要 · Abstract (English)
Recent work in machine learning increasingly attributes human-like capabilities such as reasoning or theory of mind to large language models (LLMs) on the basis of benchmark performance. This paper examines this practice through the lens of construct validity, understood as the problem of linking theoretical capabilities to their empirical measurements. It contrasts three influential frameworks: the nomological account developed by Cronbach and Meehl, the inferential account proposed by Messick and refined by Kane, and Borsboom's causal account. I argue that the nomological account provides the most suitable foundation for current LLM capability research. It avoids the strong ontological commitments of the causal account while offering a more substantive framework for articulating construct meaning than the inferential account. I explore the conceptual implications of adopting the nomological account for LLM research through a concrete case: the assessment of reasoning capabilities in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。