arXiv:2602.23949cs.IRcs.AI2026-02Conference of the …

评测智能搜索系统质量与效率的平衡,发现大模型存在冗余调用问题。

HotelQuEST: Balancing Quality and Efficiency in Agentic Search

  • 构建214条酒店查询数据集,覆盖从简单到复杂的全难度范围
  • 发现大模型代理准确率高但耗时耗能,主因是重复调用和路径错配
  • 提供明确偏好标注,适合评估真实场景下用户意图不完整的问题

智能搜索作为由大语言模型驱动的自适应检索系统的新范式正迅速发展。然而,现有基准主要关注性能质量,忽视了实际部署中至关重要的效率因素。此外,真实用户查询常包含未明确表达的需求偏好,这一挑战在当前评估中仍被严重低估。因此,尽管许多智能搜索系统表现优异,却难以落地应用。本文提出HotelQuEST基准,包含214条酒店搜索查询,涵盖从简单事实请求到复杂多步查询的全谱系难度。为解决偏好不明确问题,我们收集了标注者隐含偏好的澄清信息,使评估更具可解释性。实验发现,基于大模型的代理系统虽比传统检索器更准确,但因冗余工具调用和低效路由策略导致成本显著上升,未能根据查询复杂度匹配模型能力。分析揭示了当前系统的效率瓶颈,并展现出显著的成本优化潜力。

原文摘要 · Abstract (English)

Agentic search has emerged as a promising paradigm for adaptive retrieval systems powered by large language models (LLMs). However, existing benchmarks primarily focus on quality, overlooking efficiency factors that are critical for real-world deployment. Moreover, real-world user queries often contain underspecified preferences, a challenge that remains largely underexplored in current agentic search evaluation. As a result, many agentic search systems remain impractical despite their impressive performance. In this work, we introduce HotelQuEST, a benchmark comprising 214 hotel search queries that range from simple factual requests to complex queries, enabling evaluation across the full spectrum of query difficulty. We further address the challenge of evaluating underspecified user preferences by collecting clarifications that make annotators' implicit preferences explicit for evaluation. We find that LLM-based agents achieve higher accuracy than traditional retrievers, but at substantially higher costs due to redundant tool calls and suboptimal routing that fails to match query complexity to model capability. Our analysis exposes inefficiencies in current agentic search systems and demonstrates substantial potential for cost-aware optimization.

智能搜索大模型效率优化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。