用二维评估框架解决大模型部署中稳定性难题
CreditAudit: 2$^\text{nd}$ Dimension for LLM Evaluation and Selection
- 在多种提示模板下测试模型,评估平均表现与波动性
- 相同均值表现的模型波动差异可达30%以上
- 生成信用评级,帮助高风险场景选型
公共基准上的排行榜分数持续上升并趋于收敛,许多前沿语言模型之间的差距仅在微小范围内。然而这些分数往往无法反映用户日常使用体验,因为系统提示、输出协议和交互模式在迭代中不断演变,而在代理式多步骤流程中,微小的协议变化可能引发严重故障,使实践者难以抉择部署模型。我们提出CreditAudit,一种面向部署的信用审计框架,在多个基准上采用语义对齐且非对抗性的系统提示模板族评估模型,报告平均能力(跨场景平均表现)和场景波动标准差(作为稳定性风险信号),并通过跨模型分位数将波动映射为可解释的信用等级(AAA至BBB),辅以诊断工具缓解模板难度漂移。在GPQA、TruthfulQA和MMLU Pro上的受控实验表明,具有相似平均能力的模型可能表现出显著不同的波动性,稳定性风险可在代理或高失败成本场景中颠覆优先级决策。通过提供基于2D指标和等级的语言,CreditAudit支持分层部署和更严谨的测试监控资源配置,为真实世界应用提供更客观可信的模型评估。
原文摘要 · Abstract (English)
Leaderboard scores on public benchmarks have been steadily rising and converging, with many frontier language models now separated by only marginal differences. However, these scores often fail to match users' day to day experience, because system prompts, output protocols, and interaction modes evolve under routine iteration, and in agentic multi step pipelines small protocol shifts can trigger disproportionate failures, leaving practitioners uncertain about which model to deploy. We propose CreditAudit, a deployment oriented credit audit framework that evaluates models under a family of semantically aligned and non adversarial system prompt templates across multiple benchmarks, reporting mean ability as average performance across scenarios and scenario induced fluctuation sigma as a stability risk signal, and further mapping volatility into interpretable credit grades from AAA to BBB via cross model quantiles with diagnostics that mitigate template difficulty drift. Controlled experiments on GPQA, TruthfulQA, and MMLU Pro show that models with similar mean ability can exhibit substantially different fluctuation, and stability risk can overturn prioritization decisions in agentic or high failure cost regimes. By providing a 2D and grade based language for regime specific selection, CreditAudit supports tiered deployment and more disciplined allocation of testing and monitoring effort, enabling more objective and trustworthy model evaluation for real world use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。