arXiv:2505.20738cs.CL2025-05NeurIPS被引 3

提出消除大模型自生成测评集偏见的框架Silencer,提升评测可靠性。

Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-Generator

  • 利用多生成器在样本和评测集层面的差异来中和偏见
  • 将自生成评测的性能偏差降至接近零,相关性提升至0.833
  • 适用于多种场景,适合关注模型评测公平性的研究者

LLM-as-Benchmark-Generator 方法被广泛用于替代人工标注以实现可扩展评估,但该范式中的潜在偏见仍缺乏深入研究。本文系统定义并验证了模型在自生成测评集上表现虚高的现象,称为自偏见(self-bias),并将其归因于问题领域、语言风格及错误标签带来的子偏见。基于此,我们提出 Silencer 框架,通过利用多个生成器在样本和评测集层面的异质性,有效中和偏见,生成高质量且无自偏见的测评集。实验结果表明,Silencer 可将自偏见抑制至接近零,在多种设置下显著提升生成测评集的评估有效性(与高质量人工标注基准的皮尔逊相关系数从0.655提升至0.833),同时展现出强泛化能力。

原文摘要 · Abstract (English)

LLM-as-Benchmark-Generator methods have been widely studied as a supplement to human annotators for scalable evaluation, while the potential biases within this paradigm remain underexplored. In this work, we systematically define and validate the phenomenon of inflated performance in models evaluated on their self-generated benchmarks, referred to as self-bias, and attribute it to sub-biases arising from question domain, language style, and wrong labels. On this basis, we propose Silencer, a general framework that leverages the heterogeneity between multiple generators at both the sample and benchmark levels to neutralize bias and generate high-quality, self-bias-silenced benchmark. Experimental results across various settings demonstrate that Silencer can suppress self-bias to near zero, significantly improve evaluation effectiveness of the generated benchmark (with an average improvement from 0.655 to 0.833 in Pearson correlation with high-quality human-annotated benchmark), while also exhibiting strong generalizability.

大模型评测自偏见基准生成偏见消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。