arXiv:2508.02926cs.LGcs.AI2025-08

为动态需求设计可追溯的多评委模型评估协议

GrandJury: A Collaborative Machine Learning Model Evaluation Protocol for Dynamic Quality Rubrics

  • 引入时间衰减聚合与透明评分规则,支持动态评估
  • 通过多人评审捕捉共识与分歧,避免单一指标误导
  • 适合需要适应真实场景变化的AI系统开发者

生成式机器学习模型已成为现代系统的核心,广泛应用于创作写作、摘要生成、多跳推理和情境化对话。这些模型支撑大规模AI助手、工作流自动化和自主决策。然而,可接受的输出往往并非绝对或静态,而是多元且高度依赖上下文。现有评估体系仍依赖静态基准测试,导致优化偏向排行榜分数而非实际用户需求或不断变化的现实。GrandJury提出一种正式评估协议,融合时间衰减聚合、完整可追溯性、动态透明的任务评分规则及多评委人类判断。这些机制共同实现多元、可问责的评估,捕捉动态共识并揭示分歧。我们提供了开源实现(grandjury PyPI 包)和公开的大型语言模型(LLM)推理输出数据集,以说明该方法的必要性与可行性。GrandJury为无绝对真值时的机器学习输出评估提供了新范式。

原文摘要 · Abstract (English)

Generative Machine Learning models have become central to modern systems, powering applications in creative writing, summarization, multi-hop reasoning, and context-aware dialogue. These models underpin large-scale AI assistants, workflow automation, and autonomous decision-making. In such domains, acceptable response is rarely absolute or static, but plural and highly context-dependent. Yet standard evaluation regimes still rely on static, benchmark-style tests, incentivizing optimization toward leaderboard scores rather than alignment with dynamic user needs or evolving realities. GrandJury introduces a formal evaluation protocol combining time-decayed aggregation, complete traceability, with the support of dynamic, transparent task rubric attribution, and multi-rater human judgment. Together, these elements enable pluralistic, accountable evaluation that captures evolving consensus and surfaces disagreement. We provide an open-source implementation (grandjury PyPI package) and a public collection of Large Language Model (LLM) inference outputs to illustrate the need and method. GrandJury provides a new paradigm for AI practitioners when evaluating machine learning outputs without absolute ground truth.

模型评估动态评分多评委LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。