arXiv:2603.18173cs.CL2026-03

构建持续评估框架,检测大模型随时间演进中的性能退化问题

GRAFITE: Generative Regression Analysis Framework for Issue Tracking and Evaluation

  • 基于用户反馈构建模型问题库,实现长期问题追踪
  • 利用大模型自评机制进行质量保证测试,支持多版本对比
  • 适合关注模型可靠性与迭代评估的研究者和开发者

大型语言模型(LLMs)的性能通常在发布时依赖热门话题和基准测试。然而,随着训练数据对基准集的广泛暴露,长期存在数据污染风险,可能导致评估结果虚高。为应对这一挑战,我们提出GRAFITE——一个持续的大型语言模型评估平台,通过全面系统维护与评估模型问题。该方法基于用户反馈建立模型问题仓库,并提供一套利用大模型作为裁判(LLM-as-a-judge)的质量保证测试流程,支持多模型版本并行比较,有效检测不同版本间的性能回归。平台开源地址:https://github.com/IBM/grafite,演示视频:www.youtube.com/watch?v=XFZyoleN56k。

原文摘要 · Abstract (English)

Large language models (LLMs) are largely motivated by their performance on popular topics and benchmarks at the time of their release. However, over time, contamination occurs due to significant exposure of benchmark data during training. This poses a risk of model performance inflation if testing is not carefully executed. To address this challenge, we present GRAFITE, a continuous LLM evaluation platform through a comprehensive system for maintaining and evaluating model issues. Our approach enables building a repository of model problems based on user feedback over time and offers a pipeline for assessing LLMs against these issues through quality assurance (QA) tests using LLM-as-a-judge. The platform enables side-by-side comparison of multiple models, facilitating regression detection across different releases. The platform is available at https://github.com/IBM/grafite. The demo video is available at www.youtube.com/watch?v=XFZyoleN56k.

大模型评估性能监控持续测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。