arXiv:2509.04676cs.AIcs.HC2025-09

用人类认知能力优化AI评估,让评测更贴近真实需求。

An Approach to Grounding AI Model Evaluations in Human-derived Criteria

  • 基于人类访谈与调查,提炼出优先级、记忆、辨识、情境化等关键认知技能。
  • 发现人类对AI的性能期望高,但认为其缺乏解释力和共情能力。
  • 提出可落地的评估框架,适合关注AI人性化设计的研究者使用。

在快速发展的人工智能领域,传统基准测试难以捕捉模型的细微能力。本文聚焦物理世界建模任务,提出一种新方法:将人类衍生的评估标准融入现有基准,以提升模型行为的可解释性与实用性。研究基于感知测试(Perception Test)和OpenEQA基准,通过深度访谈与大规模调查,识别出优先级判断、记忆、辨识与情境化等关键认知技能,这些技能对人类与AI推理均至关重要。结果显示,参与者普遍认为当前AI在解释与共情能力上不足,但对性能抱有较高期待。通过将这些发现整合进基准设计,本文提出一套面向人类认知过程的评估框架,为研究人员和实践者提供可操作的指导,推动AI评估向用户中心范式演进,同时为未来评估体系发展奠定基础。

原文摘要 · Abstract (English)

In the rapidly evolving field of artificial intelligence (AI), traditional benchmarks can fall short in attempting to capture the nuanced capabilities of AI models. We focus on the case of physical world modeling and propose a novel approach to augment existing benchmarks with human-derived evaluation criteria, aiming to enhance the interpretability and applicability of model behaviors. Grounding our study in the Perception Test and OpenEQA benchmarks, we conducted in-depth interviews and large-scale surveys to identify key cognitive skills, such as Prioritization, Memorizing, Discerning, and Contextualizing, that are critical for both AI and human reasoning. Our findings reveal that participants perceive AI as lacking in interpretive and empathetic skills yet hold high expectations for AI performance. By integrating insights from our findings into benchmark design, we offer a framework for developing more human-aligned means of defining and measuring progress. This work underscores the importance of user-centered evaluation in AI development, providing actionable guidelines for researchers and practitioners aiming to align AI capabilities with human cognitive processes. Our approach both enhances current benchmarking practices and sets the stage for future advancements in AI model evaluation.

AI评估认知技能人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。