首个面向个性化深度研究的基准测试,填补了AI研究员评估空白。
Towards Personalized Deep Research: Benchmarks and Evaluations
- 构建50个任务+25个用户画像的250组真实查询对,支持个性化研究评估
- 提出PQR框架,从个性契合度、内容质量、事实可靠性三方面综合评测
- 首次系统揭示现有AI研究助手在个性化场景下的能力边界
深度研究智能体(DRAs)可自主开展复杂调查并生成综合性报告,具备强大的现实应用潜力。然而,现有评估多依赖封闭式基准,开放式的深度研究基准仍稀缺,且普遍忽略个性化场景。为此,我们提出首个用于评估个性化深度研究的基准——个人化深度研究基准(PDR-Bench)。该基准将10个领域的50个多样化研究任务与25个包含结构化人格特征和动态现实背景的真实用户档案配对,生成250组贴近实际的用户-任务查询。为评估系统性能,我们设计了PQR评估框架,联合衡量个性化契合度、内容质量与事实可靠性。对多种系统的实验表明,当前系统在处理个性化深度研究任务时存在明显的能力与局限。本工作为下一代真正个性化的AI研究助手的研发与评估奠定了严谨基础。
原文摘要 · Abstract (English)
Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmarks, while open-ended deep research benchmarks remain scarce and typically neglect personalized scenarios. To bridge this gap, we introduce Personalized Deep Research Bench (PDR-Bench), the first benchmark for evaluating personalization in DRAs. It pairs 50 diverse research tasks across 10 domains with 25 authentic user profiles that combine structured persona attributes with dynamic real-world contexts, yielding 250 realistic user-task queries. To assess system performance, we propose the PQR Evaluation Framework, which jointly measures Personalization Alignment, Content Quality, and Factual Reliability. Our experiments on a range of systems highlight current capabilities and limitations in handling personalized deep research. This work establishes a rigorous foundation for developing and evaluating the next generation of truly personalized AI research assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。