评测指令型检索模型在探索性搜索中的表现,发现其排序相关性提升但指令遵循能力不足。
Can Instructed Retrieval Models Really Support Exploration?
- 对比微调与提示的指令检索模型,评估其在特定探索场景下的表现。
- 最佳模型在排序相关性上优于传统方法,但指令响应不敏感或反直觉。
- 适合短期探索,不适合长期需精准响应指令的复杂查询任务。
探索性搜索以目标不明确和查询意图动态变化为特征。在此类场景中,能够捕捉用户意图细微差别并动态调整结果的检索模型尤为重要——指令跟随式检索模型正具备此潜力。本文通过专家标注的测试集,评估了近期针对指令检索微调的LLM及通用大模型使用高效成对排序提示(Pairwise Ranking Prompting)进行排序的表现。结果显示,最优指令检索模型在排序相关性上优于无指令感知的方法。然而,指令遵循性能(影响用户体验的关键指标)并未随相关性提升而改善,反而表现出对指令不敏感或反直觉的行为。研究表明,尽管当前指令检索模型相比无指令方法对用户有益,但在需要持续响应指令的长时探索会话中,其实际收益有限。
原文摘要 · Abstract (English)
Exploratory searches are characterized by under-specified goals and evolving query intents. In such scenarios, retrieval models that can capture user-specified nuances in query intent and adapt results accordingly are desirable -- instruction-following retrieval models promise such a capability. In this work, we evaluate instructed retrievers for the prevalent yet under-explored application of aspect-conditional seed-guided exploration using an expert-annotated test collection. We evaluate both recent LLMs fine-tuned for instructed retrieval and general-purpose LLMs prompted for ranking with the highly performant Pairwise Ranking Prompting. We find that the best instructed retrievers improve on ranking relevance compared to instruction-agnostic approaches. However, we also find that instruction following performance, crucial to the user experience of interacting with models, does not mirror ranking relevance improvements and displays insensitivity or counter-intuitive behavior to instructions. Our results indicate that while users may benefit from using current instructed retrievers over instruction-agnostic models, they may not benefit from using them for long-running exploratory sessions requiring greater sensitivity to instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。