arXiv:2509.26632cs.AI2025-09

用分层图结构整合多维度指标,让AI评估更透明可解释。

Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees

  • 构建分层有向图,将多种评估信号按用户定义方式聚合
  • 在大规模测评中验证了方法对异构证据的整合能力
  • 适合需要透明、多维度评估AI系统的研究者与工程师

本文提出测量树(measurement trees),一种新型度量方法,通过层次化的有向图将多种评估构建块整合为可解释的多级表示。与传统单值、向量或分类度量不同,测量树中的每个节点通过用户定义的聚合方式总结其子节点信息。针对当前AI评估范围需拓展的呼声,该方法提升了度量透明性,支持融合异构证据,如代理行为、商业表现、能效、社会技术因素及安全信号等。文中给出定义与实例,通过大规模测量实验展示实用性,并提供开源Python代码。该工作为复杂概念的透明化测量提供了系统性基础,推动更广泛且可解释的AI评估体系发展。

原文摘要 · Abstract (English)

This paper introduces \textit{measurement trees}, a novel class of metrics designed to combine various constructs into an interpretable multi-level representation of a measurand. Unlike conventional metrics that yield single values, vectors, surfaces, or categories, measurement trees produce a hierarchical directed graph in which each node summarizes its children through user-defined aggregation methods. In response to recent calls to expand the scope of AI system evaluation, measurement trees enhance metric transparency and facilitate the integration of heterogeneous evidence, including, e.g., agentic, business, energy-efficiency, sociotechnical, or security signals. We present definitions and examples, demonstrate practical utility through a large-scale measurement exercise, and provide accompanying open-source Python code. By operationalizing a transparent approach to measurement of complex constructs, this work offers a principled foundation for broader and more interpretable AI evaluation.

AI评估可解释性多维度指标测量框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。