用对抗性改写让评测集防记忆,保持长期有效性
LastingBench: Defend Benchmarks Against Knowledge Leakage
- 通过扰动识别数据泄露点,生成反事实内容
- 主流QA基准测试性能下降显著,证明有效抑制记忆
- 适合关注模型评估公平性的研究者和开发者
大语言模型日益复杂,存在通过记忆特定任务数据‘作弊’标准问答评测的问题,导致评测结果失真。尽管已有研究聚焦于检测泄漏,但对缓解其影响、保障评测长期可用性关注不足。本文提出LastingBench框架,持续强化现有评测集以抵御知识泄漏。该框架通过扰动识别泄漏点,并将其重写为反事实内容,在破坏记忆路径的同时保留原评测意图。对当前主流QA基准的评估显示显著性能差距,验证了其有效降低记忆效应的能力。LastingBench提供了一种实用且可扩展的方案,确保评测长期稳健,推动更公平、可解释的LLM评估。
原文摘要 · Abstract (English)
The increasing complexity of large language models (LLMs) raises concerns about their ability to "cheat" on standard Question Answering (QA) benchmarks by memorizing task-specific data. This undermines the validity of benchmark evaluations, as they no longer reflect genuine model capabilities but instead the effects of data leakage. While prior work has focused on detecting such leakage, little attention has been given to mitigating its impact and preserving the long-term utility of benchmarks. In this paper, we introduce LastingBench, a novel framework designed to continuously reinforce and safeguard existing benchmarks against knowledge leakage. LastingBench identifies leakage points in the context through perturbation, then rewrites the leakage points to counterfactual ones-disrupting memorization while preserving the benchmark's original evaluative intent. Evaluations of state-of-the-art QA benchmarks show significant performance gaps, highlighting the efficacy of LastingBench in reducing memorization effects. LastingBench offers a practical and scalable solution to ensure benchmark robustness over time, promoting fairer and more interpretable evaluations of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。