arXiv:2509.25106cs.CLcs.AI2025-09被引 13

首个面向个性化深度研究的基准测试,填补了AI研究员评估空白。

Towards Personalized Deep Research: Benchmarks and Evaluations

  • 构建50个任务+25个用户画像的250组真实查询对,支持个性化研究评估
  • 提出PQR框架,从个性契合度、内容质量、事实可靠性三方面综合评测
  • 首次系统揭示现有AI研究助手在个性化场景下的能力边界

深度研究智能体(DRAs)可自主开展复杂调查并生成综合性报告,具备强大的现实应用潜力。然而,现有评估多依赖封闭式基准,开放式的深度研究基准仍稀缺,且普遍忽略个性化场景。为此,我们提出首个用于评估个性化深度研究的基准——个人化深度研究基准(PDR-Bench)。该基准将10个领域的50个多样化研究任务与25个包含结构化人格特征和动态现实背景的真实用户档案配对,生成250组贴近实际的用户-任务查询。为评估系统性能,我们设计了PQR评估框架,联合衡量个性化契合度、内容质量与事实可靠性。对多种系统的实验表明,当前系统在处理个性化深度研究任务时存在明显的能力与局限。本工作为下一代真正个性化的AI研究助手的研发与评估奠定了严谨基础。

原文摘要 · Abstract (English)

Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmarks, while open-ended deep research benchmarks remain scarce and typically neglect personalized scenarios. To bridge this gap, we introduce Personalized Deep Research Bench (PDR-Bench), the first benchmark for evaluating personalization in DRAs. It pairs 50 diverse research tasks across 10 domains with 25 authentic user profiles that combine structured persona attributes with dynamic real-world contexts, yielding 250 realistic user-task queries. To assess system performance, we propose the PQR Evaluation Framework, which jointly measures Personalization Alignment, Content Quality, and Factual Reliability. Our experiments on a range of systems highlight current capabilities and limitations in handling personalized deep research. This work establishes a rigorous foundation for developing and evaluating the next generation of truly personalized AI research assistants.

深度研究个性化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。