arXiv:2605.07986cs.HCcs.AI2026-05

用专家访谈+AI生成,把模糊的AI应用变成可复现的评测场景。

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios

论文配图:Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
图 1 · 摘自论文原文
  • 通过结构化问卷收集领域专家的高阶应用需求
  • 结合大模型与人工审核,从6个场景生成107个可测评估场景
  • 强调真实业务落地与人类需求,适合做可信AI评估的研究者

AI评估方法多样,常导致比较结果如'苹果对橙子'。为实现真实世界中'苹果对苹果'的公平比较,本文倡导评估过程的方法透明、操作落地与以人为本设计。提出一种可重复流程:通过包含六个要素(应用场景、行业、用户、预期目标、正负影响、关键指标)的结构化问卷,从金融服务业专家处获取高阶使用案例。示例包括网络安全赋能、开发者效率提升、金融犯罪聚合、可疑活动报告(SAR)提交、信用凭证生成和内部客服支持。基于这些案例,采用三阶段扩展管道——结合大语言模型提示与人工评审——生成107个详细评估场景。每一步均设人工校验点,确保场景标题、核心要素(用户、收益与风险、度量标准)及叙事逻辑符合实际业务需求。提出评估场景质量验证评分体系,明确关键构成,推动更一致、更有意义的人本化AI评估范式。

原文摘要 · Abstract (English)

AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be "apples-to-oranges" comparisons across AI evaluations. To move toward "apples-to-apples" comparisons in real-world AI evaluations, this work advocates for methodological transparency in evaluation scenarios, operational grounding, and human-centered design (HCD) principles. We propose a repeatable process for transforming high-level use cases to detailed scenarios by eliciting use cases from subject matter experts (SMEs) via a structured AI Use Case Worksheet with six key elements: use case, sector, user (direct and indirect), intended outcomes, expected impacts (positive and negative), and KPIs and metrics. We demonstrate utility of the worksheet and process in the U.S. financial services sector. This paper reports on example high-level AI use cases identified by financial services sector SMEs: cyber defense enablement, developer productivity, financial crime aggregation, suspicious activity report (SAR) filing, credit memo generation, and internal call center support. These AI use cases provided are illustrative of the process and not exhaustive. Central to our work is a three-stage expansion pipeline combining LLM prompting with human reviews to generate 107 scenarios from those use cases elicited from SMEs. This process integrates iterative human reviews at every juncture to ensure operational grounding: for scenario titles and descriptions; for core scenario elements like users, benefits and risks, and metrics; and for scenario narratives and evaluation objectives. Human checkpoints ensure scenarios remain reflective of real-world usage and human needs. We describe a validation rubric to assess scenario quality. By defining key scenario components, this work supports a more consistent and meaningful paradigm for human-centered AI evaluations.

AI评估人本设计场景构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。