用大模型生成细粒度用户意图,自动评估搜索结果满足度。
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
- 基于用户属性和意图类型生成多样化细粒度搜索目标
- 意图级评估与专家一致率达72%,可规模化分析
- 适合搜索系统优化、用户体验研究者使用
同一搜索词可能对应100种不同目标。现有评估方法难以捕捉这种多样性。本文提出BloomIntent,将用户意图作为评估单元:首先基于用户属性与信息获取类型生成合理、细粒度的意图;再利用大语言模型自动化评估搜索结果对每种意图的满足程度。为支持实际分析,BloomIntent对语义相似意图聚类,并以结构化界面汇总评估结果。三项技术评估显示,BloomIntent生成的意图真实、可评估且具备可扩展性,意图级满意度评分与专家评价达成72%一致性。在一项案例研究(N=4)中,它帮助搜索专家识别模糊查询背后的意图,发现未被满足的需求,并提炼出改进搜索体验的具体洞察。通过从查询级转向意图级评估,BloomIntent重新定义了搜索系统的评价方式——不仅看性能,更看能否服务多元用户目标。
原文摘要 · Abstract (English)
If 100 people issue the same search query, they may have 100 different goals. While existing work on user-centric AI evaluation highlights the importance of aligning systems with fine-grained user intents, current search evaluation methods struggle to represent and assess this diversity. We introduce BloomIntent, a user-centric search evaluation method that uses user intents as the evaluation unit. BloomIntent first generates a set of plausible, fine-grained search intents grounded on taxonomies of user attributes and information-seeking intent types. Then, BloomIntent provides an automated evaluation of search results against each intent powered by large language models. To support practical analysis, BloomIntent clusters semantically similar intents and summarizes evaluation outcomes in a structured interface. With three technical evaluations, we showed that BloomIntent generated fine-grained, evaluable, and realistic intents and produced scalable assessments of intent-level satisfaction that achieved 72% agreement with expert evaluators. In a case study (N=4), we showed that BloomIntent supported search specialists in identifying intents for ambiguous queries, uncovering underserved user needs, and discovering actionable insights for improving search experiences. By shifting from query-level to intent-level evaluation, BloomIntent reimagines how search systems can be assessed -- not only for performance but for their ability to serve a multitude of user goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。