评测检索模型对用户指令的响应能力,发现多数模型仍不达标。
Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models
- 构建六维评估基准InfoSearch,覆盖受众、格式等文档属性
- 提出SICR与WISE新指标,精准衡量模型指令遵循度
- 实测显示大模型+微调仍难满足复杂指令需求
大型语言模型(LLM)的指令跟随能力显著提升,使用户能通过详细提示进行复杂交互。然而,检索系统进展滞后,仍主要依赖传统的词汇和语义匹配技术,难以全面捕捉用户意图。近期虽出现指令感知的检索模型,但大多仅关注内容相关性,忽视了文档层面属性的定制化偏好。本研究在内容相关性之外,评估各类检索模型的指令跟随能力,涵盖基于LLM的密集检索与重排序模型。我们构建了InfoSearch这一新型检索评估基准,覆盖受众、关键词、格式、语言、长度和来源共六种文档级属性,并引入严格指令合规率(SICR)和加权指令敏感性评估(WISE)两项新指标,以准确衡量模型对指令的响应能力。研究结果表明,尽管在指令感知数据集上微调及增大模型规模可提升性能,但多数模型仍未能达到理想的指令合规水平。
原文摘要 · Abstract (English)
Instruction-following capabilities in LLMs have progressed significantly, enabling more complex user interactions through detailed prompts. However, retrieval systems have not matched these advances, most of them still relies on traditional lexical and semantic matching techniques that fail to fully capture user intent. Recent efforts have introduced instruction-aware retrieval models, but these primarily focus on intrinsic content relevance, which neglects the importance of customized preferences for broader document-level attributes. This study evaluates the instruction-following capabilities of various retrieval models beyond content relevance, including LLM-based dense retrieval and reranking models. We develop InfoSearch, a novel retrieval evaluation benchmark spanning six document-level attributes: Audience, Keyword, Format, Language, Length, and Source, and introduce novel metrics -- Strict Instruction Compliance Ratio (SICR) and Weighted Instruction Sensitivity Evaluation (WISE) to accurately assess the models' responsiveness to instructions. Our findings indicate that although fine-tuning models on instruction-aware retrieval datasets and increasing model size enhance performance, most models still fall short of instruction compliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。