arXiv:2605.10530cs.IR2026-05中稿 · SIGIR 2026

让研究型AI根据用户水平自动调整探索深度,避免专家嫌啰嗦、新手看不懂。

Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery

论文配图:Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery
图 1 · 摘自论文原文
  • 把用户背景动态融入搜索推理流程,实时调整查资料的深浅和范围。
  • 在四个真实任务上比商业系统更准更相关,报告匹配度提升显著。
  • 适合需要个性化科研助手的研究者或教育科技开发者。

由大模型驱动的深度研究代理已实现从规划、提问到迭代网络探索的学术发现自动化,但受限于静态的‘一刀切’检索模式。现有系统无法根据用户的现有知识水平或潜在兴趣自适应调整探索的深度与广度,常导致专家看到冗余内容、新手面对信息过载。为此,我们提出个性化深度研究(PDR)框架,将动态用户上下文融入核心检索-推理循环。PDR不把个性化当作事后格式化,而是将用户画像建模与迭代查询生成、双阶段(私有/公开)检索、上下文感知合成统一整合,使系统能自主对齐研究子目标与用户意图,并优化证据收集的终止条件。为支持基准测试,我们发布涵盖四个真实用户任务的PDR数据集,并提出融合词法指标与大模型评估的混合评价框架,用于衡量事实准确性与个性化契合度。实验结果表明,相比商业基线,PDR显著提升检索效用与报告相关性,有效弥合通用信息检索与个性化知识获取之间的差距。资源已开源:https://github.com/Applied-Machine-Learning-Lab/SIGIR2026_PDR。

原文摘要 · Abstract (English)

Deep Research agents driven by LLMs have automated the scholarly discovery pipeline, from planning and query formulation to iterative web exploration. Yet they remain constrained by a static, ``one-size-fits-all'' retrieval paradigm. Current systems fail to adaptively adjust the depth and breadth of exploration based on the user's existing expertise or latent interests, frequently resulting in reports that are either redundant for experts or overly dense for novices. To address this, we introduce Personalized Deep Research (PDR), a framework that integrates dynamic user context into the core retrieval-reasoning loop. Rather than treating personalization as a post-hoc formatting step, PDR unifies user profile modeling with iterative query development, dual-stage (private/public) retrieval, and context-aware synthesis. This allows the system to autonomously align research sub-goals with user intent and optimize the stopping criteria for evidence collection. To facilitate benchmarking, we release the PDR Dataset, covering four realistic user tasks, and propose a hybrid evaluation framework combining lexical metrics with LLM-based judgments to assess factual accuracy and personalization alignment. Experimental results against commercial baselines demonstrate that PDR significantly improves retrieval utility and report relevance, effectively bridging the gap between generic information retrieval and personalized knowledge acquisition. The resource is available to the public at https://github.com/Applied-Machine-Learning-Lab/SIGIR2026_PDR.

个性化研究代理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。