用分层贝叶斯模型修正大模型评测中提示词依赖问题,提升结果可靠性。
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering
- 基于嵌入空间聚类构建分层贝叶斯模型,捕捉提示词相关性
- 在数据有限时降低误差4-73%,提升后验密度40-450单位
- 适合做模型鲁棒性评测、小样本场景下的性能评估者
大模型评测指标常因两个不成立的假设而失准:一是评估次数充足,二是测试提示词相互独立。本文提出一种结合嵌入空间聚类的分层贝叶斯修正模型,在数据有限情况下仍能提供稳健性能度量,并纠正提示词依赖问题。应用于对抗鲁棒性评测,该方法稳定恢复了提示词的聚类结构,显著提升评估可靠性,平均绝对误差降低4%至73%,预期对数后验密度提升40至450单位。
原文摘要 · Abstract (English)
LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations are available for classical inference, and (ii) test prompts are independent. We propose a corrective Bayesian hierarchical model with embedding-space clustering that provides robust performance metrics in limited-data settings while correcting for prompt dependence. We apply the approach to adversarial robustness benchmarks, showing consistent recovery of clustering structure, resulting in more reliable performance metrics, with 4-73% improvements to mean absolute errors and 40-450 unit improvements to expected log posterior densities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。