arXiv:2608.03340cs.CL2026-08

测试常识基准对真实任务的预测能力,发现其效果有限且依赖具体任务。

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

论文配图:Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks
图 1 · 摘自论文原文
  • 在23个模型上对比多个常识与下游任务,评估基准预测力。
  • 修订版基准未提升预测效果,仅少数任务有小幅增益。
  • 常识基准只对特定下游任务有效,不具备广泛适用性。

预测大模型在真实任务中的表现至关重要,但现有常识基准对下游性能的预测能力仍不明确。为评估常用常识基准的实际有效性,我们在四个经典常识基准、四个改进版本、三个非常识对照任务以及八个需要隐含社会、语用、时间或物理推理的下游任务上,对六家族共23个模型进行了评测。通过比较模型排名、控制计算量后的相关性,并采用留一家族交叉验证评估基准的效度。结果表明,修订后的基准基本保持原有模型排序,但未显著提升对下游任务的预测能力;常识基准仅在极少数下游任务中展现出跨家族一致性预测能力,其他任务中收益微小或仅限特定指标。总体而言,标准化常识基准只能提供针对具体任务的证据,而非普遍意义上的常识能力证明。

原文摘要 · Abstract (English)

Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.

常识推理模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。