让搜索代理更懂网页结构,提升推理与检索的匹配度
Rethinking Deep Research from the Perspective of Web Content Distribution Matching
- 将网页索引结构纳入观察空间,动态调整搜索策略
- 通过少量查询估计匹配分数,提升子目标完成率
- 无需修改原有框架,适配多种复杂推理任务
尽管整合了搜索工具,深度搜索代理仍面临推理型查询与底层网页索引结构不匹配的问题。现有框架将搜索引擎视为静态工具,导致查询过于粗略或过于细粒度,难以获取精确证据。我们提出 WeDas:一种感知网页内容分布的框架,将搜索空间的结构特征引入代理的观察空间。核心是查询-结果对齐评分,用于量化代理意图与检索结果的兼容性。为应对动态网页索引难以建模的问题,我们引入少样本探测机制,通过有限查询迭代估算该评分,使代理能根据局部内容环境动态重设子目标。作为即插即用模块,WeDas 在四个基准上持续提升子目标完成率与准确性,有效弥合高层推理与底层检索之间的鸿沟。
原文摘要 · Abstract (English)
Despite the integration of search tools, Deep Search Agents often suffer from a misalignment between reasoning-driven queries and the underlying web indexing structures. Existing frameworks treat the search engine as a static utility, leading to queries that are either too coarse or too granular to retrieve precise evidence. We propose WeDas, a Web Content Distribution Aware framework that incorporates search-space structural characteristics into the agent's observation space. Central to our method is the Query-Result Alignment Score, a metric quantifying the compatibility between agent intent and retrieval outcomes. To overcome the intractability of indexing the dynamic web, we introduce a few-shot probing mechanism that iteratively estimates this score via limited query accesses, allowing the agent to dynamically recalibrate sub-goals based on the local content landscape. As a plug-and-play module, WeDas consistently improves sub-goal completion and accuracy across four benchmarks, effectively bridging the gap between high-level reasoning and low-level retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。