arXiv:2510.03795cs.IRcs.CL2025-10被引 9

研究大模型在个性化对话检索中的表现波动,发现人工选知识比模型自动选更可靠。

Investigating LLM Variability in Personalized Conversational Information Retrieval

  • 用多轮实验验证不同大模型在个性化检索中的输出差异
  • 人工挑选的知识库始终优于模型自动选择,提升召回率
  • 建议采用多轮评估并报告方差,尤其关注召回类指标

个性化对话信息检索(CIR)近年来快速发展,依赖于大语言模型(LLM)的进步。该任务通过用户偏好、知识或约束等个性化信息优化文档检索。本研究基于TREC iKAT 2023数据集复现并扩展了相关工作,聚焦于LLM输出的可变性与泛化能力。我们使用新发布的TREC iKAT 2024数据集,评估包括Llama(1B-70B)、Qwen-7B、GPT-4o-mini在内的多种模型。结果表明,经人工筛选的个人文本知识库(PTKB)能持续提升检索性能,而基于LLM的筛选方法未能稳定超越人工选择。跨数据集比较显示,iKAT的变异性高于CAsT,凸显个性化CIR评估挑战。值得注意的是,以召回为导向的指标方差低于以精确率为导向的指标,这对第一阶段检索器设计具有关键意义。研究强调应开展多轮评估并报告方差,以推动更稳健、可推广的个性化CIR评估实践。

原文摘要 · Abstract (English)

Personalized Conversational Information Retrieval (CIR) has seen rapid progress in recent years, driven by the development of Large Language Models (LLMs). Personalized CIR aims to enhance document retrieval by leveraging user-specific information, such as preferences, knowledge, or constraints, to tailor responses to individual needs. A key resource for this task is the TREC iKAT 2023 dataset, designed to evaluate personalization in CIR pipelines. Building on this resource, Mo et al. explored several strategies for incorporating Personal Textual Knowledge Bases (PTKB) into LLM-based query reformulation. Their findings suggested that personalization from PTKBs could be detrimental and that human annotations were often noisy. However, these conclusions were based on single-run experiments using the GPT-3.5 Turbo model, raising concerns about output variability and repeatability. In this reproducibility study, we rigorously reproduce and extend their work, focusing on LLM output variability and model generalization. We apply the original methods to the new TREC iKAT 2024 dataset and evaluate a diverse range of models, including Llama (1B-70B), Qwen-7B, GPT-4o-mini. Our results show that human-selected PTKBs consistently enhance retrieval performance, while LLM-based selection methods do not reliably outperform manual choices. We further compare variance across datasets and observe higher variability on iKAT than on CAsT, highlighting the challenges of evaluating personalized CIR. Notably, recall-oriented metrics exhibit lower variance than precision-oriented ones, a critical insight for first-stage retrievers. Finally, we underscore the need for multi-run evaluations and variance reporting when assessing LLM-based CIR systems. By broadening evaluation across models, datasets, and metrics, our study contributes to more robust and generalizable practices for personalized CIR.

个性化检索大模型评估可重复性召回率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。