警惕大模型评估中尾部形状估计的假阳性,提出诊断协议。
Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives
- 设计可预注册的评估协议,涵盖适切性、拟合度等四重检验
- 在两种评分器下发现三种假阳性模式,原宣称的尾部形状结论被推翻
- 建议将该协议作为大模型毒性评估中尾部指数声明的基准
近期研究主张将大语言模型(LLM)评估从均值指标转向关注尾部的度量,包括条件风险价值和奖励模型误差的尾部指数估计。本文探究经典极值理论中的尾部指数参数是否在均值和尾部幅度统计之外提供额外区分信息。我们预注册了一套评估协议,包含适切性、拟合优度、阈值稳定性及效应量要求,用于验证任何正尾部形状主张。该协议是本文核心贡献;后续实证研究仅作为其检测能力的示范。在两种结构不同的评分器下应用于标准的LLM毒性评估场景,协议识别出三种被简单分析忽略的假阳性模式,并驳回了两个评分器的尾部形状宣称。结论表明,在所考察的毒性评估设置中,尾部形状估计的可靠性远低于近期文献所暗示,建议将本协议作为类似场景中尾部指数声明的起点。
原文摘要 · Abstract (English)
Recent work motivates moving large language model (LLM) evaluation from mean-based to tail-aware metrics, including conditional value-at-risk and tail-index estimates of reward-model error. We ask whether the canonical extreme-value-theory tail-index parameter, which isolates how heavy a tail is from how large the tail mass is, adds discriminative information beyond the mean and a standard tail-magnitude statistic in LLM evaluation. We pre-register a protocol covering admissibility, goodness-of-fit, threshold-stability, and effect-size requirements for any positive tail-shape claim. The protocol is the contribution of this paper; the empirical study below is a demonstration of what its gates catch. Applied to a standard LLM toxicity-evaluation setup under two structurally different scorer families, the protocol catches three distinct modes of false positives that a naive analysis would have published, and rejects the headline tail-shape claim on both scorers. We conclude that tail-shape estimation in the LLM toxicity-evaluation setups we examined is more fragile than the recent literature suggests, and recommend the protocol as a starting point for tail-index claims in similar setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。