arXiv:2605.17829cs.AI2026-05被引 1

交互式AI评估需系统化设计,不能沿用旧方法。

Interactive Evaluation Requires a Design Science

论文配图:Interactive Evaluation Requires a Design Science
图 1 · 摘自论文原文
  • 将评估视为证据到判断的映射,强调交互轨迹为新证据
  • 提出双轴分类法与设计原则,统一评估标准
  • 适合构建智能体系统的研究人员参考

AI评估正经历结构性变革。大语言模型(LLMs)越来越多地作为通过工具、环境、用户和其他智能体持续互动的系统部署,但许多评估实践仍沿用以响应为中心的基准假设(如固定输入、孤立输出,以及仅基于单次响应的评判)。尽管已开始构建交互式基准,但当前生态碎片化:不同基准在允许的交互产物、轨迹评分方式及结论支持范围上差异显著。本文主张,交互式评估应被视为一个有原则的评估范式,而非仅仅是新的智能体基准集合。简单套用既有评估范式不足以应对挑战。我们定义评估为从证据到判断的自主映射,并指出交互式评估改变了映射的两端:证据变为交互生成的轨迹,评估过程必须考察流程性、可复现性、协作性、鲁棒性及系统级表现。基于此定义,提出双轴分类法,推导设计原则与报告标准,分析代表性场景,并揭示长期存在的评估难题如何在轨迹层面重现。

原文摘要 · Abstract (English)

AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other agents, while many evaluation practices still inherit assumptions from response-centered benchmarks (e.g., fixed inputs, isolated outputs, and outcome judgments that can be made from a single response). The field has begun to build interactive benchmarks, but the resulting landscape is fragmented: benchmarks differ in what interaction artifacts they admit, how trajectories are scored, and what claims their results support. This position paper argues that interactive evaluation should be treated as a principled evaluation paradigm, not merely a new family of agent benchmarks. Simply adopting previous evaluation paradigms does not suffice. We define evaluation as an autonomous mapping from evidence to judgments, and show that interactive evaluation changes both sides of this mapping: the evidence becomes interaction-generated trajectories, while the evaluation procedure must assess process, recoverability, coordination, robustness, and system-level performance. Building on this definition, we propose a two-axis taxonomy, derive design principles and reporting standards, examine representative scenarios, and analyze how longstanding evaluation challenges reappear at the trajectory level.

AI评估交互式系统设计科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。