用认知模型暴露AI评测中的隐含假设,提升评测科学性
Exposing Assumptions in AI Benchmarks through Cognitive Modelling
- 用结构方程模型显式建模评测背后的认知假设
- 揭示跨语言对齐任务中缺失的数据集与测量缺陷
- 适合关注评测可信度与理论基础的研究者
文化AI评测常依赖对测量概念的隐含假设,导致表述模糊、效度不足且各概念间关系不清。本文提出通过显式的认知模型(以结构方程模型形式)暴露这些假设。以跨语言对齐迁移为例,展示该方法如何回答关键研究问题并识别缺失数据集。该框架为评测构建提供理论基础,指导数据集开发以改进构念测量。通过增强透明度,推动更严谨、可累积的AI评估科学发展,呼吁研究者反思其评估根基。
原文摘要 · Abstract (English)
Cultural AI benchmarks often rely on implicit assumptions about measured constructs, leading to vague formulations with poor validity and unclear interrelations. We propose exposing these assumptions using explicit cognitive models formulated as Structural Equation Models. Using cross-lingual alignment transfer as an example, we show how this approach can answer key research questions and identify missing datasets. This framework grounds benchmark construction theoretically and guides dataset development to improve construct measurement. By embracing transparency, we move towards more rigorous, cumulative AI evaluation science, challenging researchers to critically examine their assessment foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。