arXiv:2605.06856cs.LGcs.CL2026-05

生成式AI在测试集表现好,但实际用处小,需从人类成果角度评估真实价值。

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility

论文配图:Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
图 1 · 摘自论文原文
  • 提出四阶段评估框架SCU-GenEval,聚焦用户目标与长期效果
  • 发现三大评估缺陷:代理替代、时间退化、分布隐藏
  • 强调应以人类能力提升为标准,而非单纯看模型输出质量

生成式AI在28个实际部署案例中——涵盖教育、医疗、软件工程和法律领域——虽在标准基准上表现优异,却未能实现真正实用价值。我们识别出评估实践中的三大重复性失败:代理替代、时间退化与分布隐藏。现有评估主要关注模型输出属性,而部署成功取决于人机交互是否持续提升利益相关者达成目标的能力。因此,关键缺失的是‘效用’:在特定场景下,通过持续使用AI系统对用户能力产生的改变。为此,我们提出SCU-GenEval框架,包含利益相关者-目标映射、构念-指标定义、机制建模与纵向效用测量四个阶段,并引入结构化部署协议、情境化用户模拟器及角色-目标驱动的代理度量工具以支持落地。最终主张,生成式AI的进步必须以可衡量的人类成果改善为标准,而非仅依赖基准得分。

原文摘要 · Abstract (English)

Generative AI systems achieve impressive performance on standard benchmarks yet fail to deliver real-world utility, a disconnect we identify across 28 deployment cases spanning education, healthcare, software engineering, and law. We argue that this benchmark utility gap arises from three recurring failures in evaluation practice: proxy displacement, temporal collapse, and distributional concealment. Motivated by these observations, we argue that generative AI evaluation requires a paradigm shift from static benchmark-centered transparency toward stakeholder, goal, and context-conditioned utility transparency grounded in human outcome trajectories. Existing evaluations primarily characterize properties of model outputs, while deployment success depends on whether interaction with AI improves stakeholders' ability to achieve their goals over time. The missing construct is therefore utility: the change in a stakeholder's capability induced through sustained interaction with an AI system within a deployment context. To operationalize this perspective, we propose SCU-GenEval, a four-stage evaluation framework consisting of stakeholder-goal mapping, construct-indicator specification, mechanism modeling, and longitudinal utility measurement. To make these stages practically deployable, we introduce three supporting instruments: structured deployment protocols, context-conditioned user simulators, and persona- and goal-conditioned proxy metrics. We conclude with domain-specific calls to action, arguing that progress in generative AI must be evaluated through measurable improvements in human outcomes rather than benchmark performance alone.

生成式AI评估框架真实效用人类成果

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。