用大模型动态预测网页内容过期时间,让搜索结果更贴合用户意图。
RAG-Enhanced Large Language Models for Dynamic Content Expiration Prediction in Web Search

- 基于大模型分析查询意图,动态推断内容有效期限。
- 在百度搜索线上测试中显著提升搜索新鲜度与用户体验。
- 适合关注搜索排序优化和信息时效性的工程师与研究者。
在商业网络搜索中,信息时效性与用户意图的对齐仍具挑战,因信息生命周期差异极大。传统工业方法依赖静态时间窗口过滤,导致‘一刀切’排序——内容虽时间新但语义已过时。为此,我们提出一种基于大语言模型(LLMs)的查询感知动态内容过期预测框架,部署于百度搜索系统,将时效性重定义为动态有效性推理任务。该框架从文档中提取细粒度时间上下文,利用大模型推断与查询相关的‘有效期限边界’——即信息因用户意图而失效的语义阈值。结合强鲁棒性幻觉抑制策略保障可靠性,该方法通过离线及在线A/B测试在真实生产流量上验证。结果表明,在搜索新鲜度与用户体验指标上均有显著提升,验证了大模型推理在工业级语义过期问题上的有效性。
原文摘要 · Abstract (English)
In commercial web search, aligning content freshness with user intent remains challenging due to the highly varied lifespans of information. Traditional industrial approaches rely on static time-window filtering, resulting in "one-size-fits-all" rankings where content may be chronologically recent but semantically expired. To address the limitation, we present a novel Large Language Models (LLMs)-based Query-Aware Dynamic Content Expiration Prediction Framework deployed in Baidu search, reformulating timeliness as a dynamic validity inference task. Our framework extracts fine-grained temporal contexts from documents and leverages LLMs to deduce a query-specific "validity horizon"-a semantic boundary defining when information becomes obsolete based on user intent. Integrated with robust hallucination mitigation strategies to ensure reliability, our approach has been evaluated through offline and online A/B testing on live production traffic. Results demonstrate significant improvements in search freshness and user experience metrics, validating the effectiveness of LLM-driven reasoning for solving semantic expiration at an industrial scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。