用真实浏览记录评测个性化网页代理,提升任务理解能力。
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

- 基于真实浏览轨迹构建个性化代理评测基准
- 新框架在多任务上显著优于已有记忆方法
- 适合研究个性化智能体与用户行为建模的学者
大语言模型的发展使网页代理能够自主执行复杂任务。现实中用户常提供不完整指令,需代理从原始浏览历史中推断上下文。现有评测基准未能捕捉这种个性化需求,或仅限于完全明确的提示,或把浏览历史简化为抽象形式。为此,我们提出PersonaTrail,一个在受控开放网络环境下评估个性化网页代理的基准。通过利用真实的浏览轨迹作为用户历史,PersonaTrail评估代理从过往会话中推断偏好和回忆信息的能力。我们进一步提出偏好感知上下文记忆(PACMem)框架,将原始浏览历史分解为两类结构化记忆:总结单次会话的事实记忆,以及提炼重复行为模式的偏好记忆。推理时,代理从这些记忆中检索最相关条目以指导个性化导航。大量实验表明,PACMem在多项任务上持续优于现有的基于记忆的基线方法。
原文摘要 · Abstract (English)
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open web environment. By leveraging realistic browsing trajectories as user history, PersonaTrail evaluates an agent's ability to infer user preferences and recall information from past browsing sessions. We further propose Preference-Aware Contextual Memory (PACMem), a framework that decomposes raw browsing histories into two types of structured memory: factual memories that summarize individual sessions and preference memories that distill recurring behavioral patterns. At inference time, the agent retrieves the most relevant entries from these memories to guide personalized navigation. Extensive experiments show that PACMem consistently outperforms existing memory-based baselines on both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。