构建大规模真实搜索数据集,支撑更可信的联邦排序学习研究
A Large-Scale Web Search Dataset for Federated Online Learning to Rank
- 基于260万条真实用户查询与点击数据,支持真实用户划分
- 包含用户标识和时间戳,可模拟异步联邦学习场景
- 解决传统基准依赖模拟点击和同步训练的局限性
集中式收集搜索交互日志用于训练排序模型引发严重隐私问题。联邦在线学习排序(FOLTR)通过不共享原始用户数据实现协作建模,提供隐私保护方案。然而现有FOLTR基准大多基于经典排序数据集的随机划分、模拟用户点击及同步客户端参与假设,过度简化真实场景,削弱实验结果可靠性。本文提出AOL4FOLTR,一个包含260万条查询、来自10,000名用户的大型网络搜索数据集。该数据集通过引入用户标识、真实点击数据与查询时间戳,克服了现有基准的关键缺陷,支持真实用户划分、行为建模及异步联邦学习设置。
原文摘要 · Abstract (English)
The centralized collection of search interaction logs for training ranking models raises significant privacy concerns. Federated Online Learning to Rank (FOLTR) offers a privacy-preserving alternative by enabling collaborative model training without sharing raw user data. However, benchmarks in FOLTR are largely based on random partitioning of classical learning-to-rank datasets, simulated user clicks, and the assumption of synchronous client participation. This oversimplifies real-world dynamics and undermines the realism of experimental results. We present AOL4FOLTR, a large-scale web search dataset with 2.6 million queries from 10,000 users. Our dataset addresses key limitations of existing benchmarks by including user identifiers, real click data, and query timestamps, enabling realistic user partitioning, behavior modeling, and asynchronous federated learning scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。