arXiv:2602.12356cs.AI2026-02

提出动态自适应评估框架,让评价更贴合真实场景与多方需求。

A Theoretical Framework for Adaptive Utility-Weighted Benchmarking

  • 用加权交互网络连接评估指标、模型组件与利益相关方
  • 引入人类参与的更新机制,使评估随需求动态演化
  • 适合关注公平性与可解释性的模型评估研究者

基准测试长期以来是机器学习乃至大型语言模型等现代AI系统的核心实践,共享任务、度量标准和排行榜为衡量进展和比较方法提供了共同基础。然而,随着AI系统在更复杂且影响深远的场景中部署,仅依赖传统基准已显不足。本文提出一个理论框架,将基准测试重构为多层自适应网络,通过加权交互连接评估指标、模型组件与利益相关方。结合联合推导的效用函数与人机协同更新规则,形式化表达了如何将人类权衡嵌入基准结构,并实现基准的动态演进,同时保持稳定性和可解释性。该框架将经典排行榜视为特例,为构建更具情境感知能力的评估协议奠定基础,可分析基准结构特性,推动更负责任、更契合人类价值观的评估体系发展。

原文摘要 · Abstract (English)

Benchmarking has long served as a foundational practice in machine learning and, increasingly, in modern AI systems such as large language models, where shared tasks, metrics, and leaderboards offer a common basis for measuring progress and comparing approaches. As AI systems are deployed in more varied and consequential settings, though, there is growing value in complementing these established practices with a more holistic conceptualization of what evaluation should represent. Of note, recognizing the sociotechnical contexts in which these systems operate invites an opportunity for a deeper view of how multiple stakeholders and their unique priorities might inform what we consider meaningful or desirable model behavior. This paper introduces a theoretical framework that reconceptualizes benchmarking as a multilayer, adaptive network linking evaluation metrics, model components, and stakeholder groups through weighted interactions. Using conjoint-derived utilities and a human-in-the-loop update rule, we formalize how human tradeoffs can be embedded into benchmark structure and how benchmarks can evolve dynamically while preserving stability and interpretability. The resulting formulation generalizes classical leaderboards as a special case and provides a foundation for building evaluation protocols that are more context aware, resulting in new robust tools for analyzing the structural properties of benchmarks, which opens a path toward more accountable and human-aligned evaluation.

评估框架人机协同动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。