揭示大模型评测中成绩虚高风险,提出审计方法提升可信度
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
- 构建路由-执行框架,通过污染测试检验评测分数可靠性
- 多模型测试发现噪声条件下分数普遍高于基准线,暗示记忆残留
- 建议在评测中加入污染敏感性与得分置信度审计,增强透明性
公共评测正日益决定大语言模型的排名、选型与部署。本文将这一以评测为中心的机制称为「硅工业官僚制」与「人工智能应试教育」,并指出其脆弱假设:评测分数直接反映真实泛化能力。然而实践中,分数可能混淆应试技巧与真正能力,尤其当训练数据中存在难以排除的污染与语义泄露时。为此,本文提出一种审计框架,用于分析大模型评测中的污染敏感性与得分置信度。采用路由-工作者架构,对比干净对照组与系统性删除、重写、扰动评测题后的噪声条件。对于真正干净的评测,噪声条件不应持续优于干净基线。然而在多个模型中,均观察到噪声条件下普遍存在且异质的高于基线的性能提升,表明评测相关线索可被重构并重新激活污染记忆。这说明相似分数背后可能蕴含显著不同的置信水平。我们主张不应否定评测本身,而应在评估中补充对污染敏感性与得分置信度的显式审计。
原文摘要 · Abstract (English)
Public benchmarks increasingly govern how large language models (LLMs) are ranked, selected, and deployed. We frame this benchmark-centered regime as Silicon Bureaucracy and AI Test-Oriented Education, and argue that it rests on a fragile assumption: that benchmark scores directly reflect genuine generalization. In practice, however, such scores may conflate exam-oriented competence with principled capability, especially when contamination and semantic leakage are difficult to exclude from modern training pipelines. We therefore propose an audit framework for analyzing contamination sensitivity and score confidence in LLM benchmarks. Using a router-worker setup, we compare a clean-control condition with noisy conditions in which benchmark problems are systematically deleted, rewritten, and perturbed before being passed downstream. For a genuinely clean benchmark, noisy conditions should not consistently outperform the clean-control baseline. Yet across multiple models, we find widespread but heterogeneous above-baseline gains under noisy conditions, indicating that benchmark-related cues may be reassembled and can reactivate contamination-related memory. These results suggest that similar benchmark scores may carry substantially different levels of confidence. Rather than rejecting benchmarks altogether, we argue that benchmark-based evaluation should be supplemented with explicit audits of contamination sensitivity and score confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。