arXiv:2502.15620cs.AIcs.LG2025-02IJCAI被引 13

梳理AI评估的六大范式,助你理解不同研究路径的差异与联系。

Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

  • 提出六种主流AI评估范式,分析其目标、方法与文化特征。
  • 揭示各范式间术语冲突与孤立发展问题,导致沟通障碍。
  • 适合关注AI评测标准、跨领域协作的研究者参考。

AI评估研究日益复杂且跨学科,吸引了来自不同背景的研究者,催生出多种评估范式。这些范式常彼此孤立发展,使用冲突的术语,忽视相互贡献,导致研究路径封闭和交流障碍,加剧了部署型AI系统未能满足预期的问题。为缓解这种封闭性,本文综述近期AI评估领域的研究成果,识别出六种主要范式,并从目标、方法与研究文化等维度,刻画各范式中的关键贡献。通过明确每种范式独特的问题与方法组合,旨在提升对当前评估方法多样性的认知,促进范式间的交叉融合。同时,指出领域内潜在空白,以启发未来研究方向。

原文摘要 · Abstract (English)

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other's contributions. This fragmentation has led to insular research trajectories and communication barriers both among different paradigms and with the general public, contributing to unmet expectations for deployed AI systems. To help bridge this insularity, in this paper we survey recent work in the AI evaluation landscape and identify six main paradigms. We characterise major recent contributions within each paradigm across key dimensions related to their goals, methodologies and research cultures. By clarifying the unique combination of questions and approaches associated with each paradigm, we aim to increase awareness of the breadth of current evaluation approaches and foster cross-pollination between different paradigms. We also identify potential gaps in the field to inspire future research directions.

AI评估研究范式跨学科综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。