搜索引擎日期筛选存在严重时间泄漏,影响回溯评估可信度。
Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study
- 用谷歌和鸭鸭搜的日期过滤器做回溯评估,发现多数结果含截止日后内容。
- 41%~55%的问题答案直接泄露,模型准确率虚高(Brier得分0.10)。
- 适合关注评估可靠性、机器学习可复现性的研究者参考。
搜索引擎的日期过滤功能被广泛用于回溯评估中,以确保检索结果在截止时间前。本文通过审计谷歌的before:过滤器和鸭鸭搜的日期范围筛选器,发现71%的谷歌问题和81%的鸭鸭搜问题中至少有一条检索结果包含截止时间后的信息,其中41%和55%的问题答案被直接揭示。使用gpt-oss-120b基于这些含泄漏文档进行预测,显示预测准确率显著虚高(Brier分数0.10对比无泄漏时的0.24)。我们识别出反复出现的泄漏机制,包括更新文章、相关推荐模块、不可靠元数据以及“缺失即暗示”信号。因此,当前主流引擎的日期限制检索不足以支撑可信的回溯评估,建议采用更强的检索防护或基于时间冻结的网页快照进行评估。
原文摘要 · Abstract (English)
Search-engine date filters are widely used to enforce pre-cutoff retrieval in retrospective evaluations of search-augmented forecasters. We show this approach is unreliable across two major search engines: auditing Google Search's before: filter and DuckDuckGo's date-range filter, we find that at least one retrieved page contains major post-cutoff leakage for 71% of questions on Google and 81% on DuckDuckGo, and the answer is directly revealed for 41% and 55%, respectively. Using gpt-oss-120b to forecast with these leaky documents, we demonstrate inflated prediction accuracy (Brier score 0.10 vs. 0.24 with leak-free documents). We characterize recurring leakage mechanisms, including updated articles, related-content modules, unreliable metadata, and absence-based signals, and argue that date-restricted search on these engines is insufficient for credible retrospective evaluation. We recommend stronger retrieval safeguards or evaluation on frozen, time-stamped web snapshots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。