别用人类测试评AI,该建专为AI设计的评估体系
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
- 主张用专为AI设计的评估框架替代人类心理测试
- 指出当前基准测试存在效度不足、数据污染等严重问题
- 适合关注AI评估方法论的研究者与开发者
大型语言模型(LLMs)在原本用于评估人类认知与心理特质的标准测试中取得了显著成绩,如智力和人格测试。尽管这些结果常被解读为模型具备类人特征,本文认为这种解读属于本体论错误。人类心理与教育测试是基于理论、针对特定人类群体校准的测量工具。未经实证验证便将这些测试应用于非人类主体,可能误判测量内容。此外,越来越多的趋势将AI在基准测试中的表现视为对“智能”等特质的衡量,但此类测试普遍存在效度问题、数据污染、文化偏见及对表面提示变化的高度敏感性。因此,将基准表现解释为类人特质缺乏充分的理论与实证依据。本文提出核心立场:停止使用人类测试评估AI,转而开发有原则的AI专用测试体系。新框架可借鉴现有心理测量学构建与验证方法,或完全从零创建,以契合AI的独特语境。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intelligence and personality. While these results are often interpreted as strong evidence of human-like characteristics in LLMs, this paper argues that such interpretations constitute an ontological error. Human psychological and educational tests are theory-driven measurement instruments, calibrated to a specific human population. Applying these tests to non-human subjects without empirical validation, risks mischaracterizing what is being measured. Furthermore, a growing trend frames AI performance on benchmarks as measurements of traits such as ``intelligence'', despite known issues with validity, data contamination, cultural bias and sensitivity to superficial prompt changes. We argue that interpreting benchmark performance as measurements of human-like traits, lacks sufficient theoretical and empirical justification. This leads to our position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead. We call for the development of principled, AI-specific evaluation frameworks tailored to AI systems. Such frameworks might build on existing frameworks for constructing and validating psychometrics tests, or could be created entirely from scratch to fit the unique context of AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。