用大模型结合用户行为数据,自动评估文档有用性,更准且省钱。
Leveraging LLMs to Evaluate Usefulness of Document
- 用上下文和用户行为引导大模型,分层判断文档有用性。
- 相比人工标注,模型生成标签更准确,提升满意度预测效果。
- 适合需要低成本高质量评估的检索系统研究者使用。
传统Cranfield评价范式因相关性与用户满意度关联弱、相关性标注成本高,难以有效衡量用户满意度。为此,本文探索利用大语言模型(LLMs)生成多层级有用性标签进行评估。提出一种以用户为中心的新框架,将用户搜索上下文与行为数据融入LLM,采用受序回归启发的级联判断结构实现多层级有用性评估。研究表明,在充分提供上下文与行为信息时,LLMs能精准评估文档有用性,其生成标签优于第三方标注方法。通过消融实验分析框架关键组件影响,并将生成标签用于预测用户满意度,真实世界实验表明该方法显著提升满意度预测模型性能。
原文摘要 · Abstract (English)
The conventional Cranfield paradigm struggles to effectively capture user satisfaction due to its weak correlation between relevance and satisfaction, alongside the high costs of relevance annotation in building test collections. To tackle these issues, our research explores the potential of leveraging large language models (LLMs) to generate multilevel usefulness labels for evaluation. We introduce a new user-centric evaluation framework that integrates users' search context and behavioral data into LLMs. This framework uses a cascading judgment structure designed for multilevel usefulness assessments, drawing inspiration from ordinal regression techniques. Our study demonstrates that when well-guided with context and behavioral information, LLMs can accurately evaluate usefulness, allowing our approach to surpass third-party labeling methods. Furthermore, we conduct ablation studies to investigate the influence of key components within the framework. We also apply the labels produced by our method to predict user satisfaction, with real-world experiments indicating that these labels substantially improve the performance of satisfaction prediction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。