arXiv:2511.18889cs.CLcs.AI2025-11ACL被引 4

用真实世界知识自动更新数据,让大模型评测更公平可靠。

CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation

  • 从原始数据提取实体关系,结合GDELT数据库注入新知识。
  • 通过重构建和迭代验证,保持语义一致且避免数据污染。
  • 适合关注评测公平性与模型真实性能的研究者。

数据污染严重影响自然语言处理中大模型评估的公平性,因模型在训练时无意接触到测试数据。现有方法或修改旧数据集,或用新收集信息生成数据,但未能完全消除模型中的预存知识,也未保留原始数据的语义复杂性。为此,我们提出CoreEval——一种自动利用真实世界知识更新数据的抗污染评估策略。该方法首先从原始数据中提取实体关系,再借助GDELT数据库获取最新相关知识,将其重新语境化并整合至原数据中,经精细化重构以确保语义连贯性和任务相关性。最终采用稳健的数据反射机制,迭代验证与优化标签,保障更新后数据与原始数据的一致性。在多个更新数据集上的实验表明,CoreEval能有效缓解由数据污染导致的性能过估问题。

原文摘要 · Abstract (English)

Data contamination poses a significant challenge to the fairness of LLM evaluations in natural language processing tasks by inadvertently exposing models to test data during training. Current studies attempt to mitigate this issue by modifying existing datasets or generating new ones from freshly collected information. However, these methods fall short of ensuring contamination-resilient evaluation, as they fail to fully eliminate pre-existing knowledge from models or preserve the semantic complexity of the original datasets. To address these limitations, we propose \textbf{CoreEval}, a \textbf{Co}ntamination-\textbf{re}silient \textbf{Eval}uation strategy for automatically updating data with real-world knowledge. This approach begins by extracting entity relationships from the original data and leveraging the GDELT database to retrieve relevant, up-to-date knowledge. The retrieved knowledge is then recontextualized and integrated with the original data, which is refined and restructured to ensure semantic coherence and enhanced task relevance. Ultimately, a robust data reflection mechanism is employed to iteratively verify and refine labels, ensuring consistency between the updated and original datasets. Extensive experiments on updated datasets validate the robustness of CoreEval, demonstrating its effectiveness in mitigating performance overestimation caused by data contamination.

大模型评测数据污染知识注入自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。