arXiv:2511.05501cs.HCcs.AI2025-11被引 1

为新闻业设计更真实可信的生成式AI评测体系

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners

  • 基于新闻从业者工作场景,构建领域导向的评测方法手册
  • 发现评估需平衡任务需求、价值取向与多方利益
  • 助力记者理解评测逻辑,提升对AI系统的判断力

基准测试在科技公司传达模型能力、研究人员及公众理解生成式AI系统方面起着关键作用。然而,现有基准常因未能充分反映真实使用情境(生态有效性)或测量底层概念(建构有效性)而受到批评。本文借鉴人机交互(HCI)方法,采用以人为本的设计流程,聚焦新闻领域,邀请23名专业从业者参与研讨会,据此设计出面向领域的评估‘操作手册’。研讨结果揭示了从业者在将具体任务转化为评估指标时面临的领域特有挑战,包括如何对齐度量标准与行业价值,以及如何协调不同利益相关方的需求。通过在新闻领域实践基于设计的基准构建方法,本研究不仅提供了一套可供从业者实验的评估框架,还提出了具有情境适配性、价值一致性并能培养领域用户评估素养的AI评测设计要求。

原文摘要 · Abstract (English)

Benchmarks play a significant role in how technology companies communicate about model capabilities and how researchers and the public understand generative AI systems. However, existing benchmarks have been criticized for their failure to adequately capture real-world usages (i.e. ecological validity) or to measure underlying concepts (i.e. construct validity). Building on approaches in HCI, we adopt a human-centered design process to address such critiques. Working within the journalism domain we engaged 23 professionals in a workshop which informed the design of a domain-oriented evaluation ``cookbook''. Our workshop findings surface domain-specific challenges and tensions faced by designers in translating specific tasks to evaluation constructs, aligning metrics with domain-specific values, and balancing needs among different stakeholders when constructing evaluations. Through an instantiation of design-based approaches for benchmark creation in the journalism domain, this work not only produces an evaluation structure for journalism practitioners to experiment with, but also lays out design requirements for AI evaluations that are contextualized, value-aligned, and cultivate evaluative literacy for domain end-users.

生成式AI评测体系新闻业人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。