用模拟搜索世界实现多跳推理,让模型更懂复杂查询
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

- 构建带类型知识图谱的虚拟搜索环境,统一数据、环境与奖励
- 通过视觉-知识交叉路径采样,保留每跳实体信息并生成奖励锚点
- 无需额外奖励模型,直接利用步骤级锚点进行策略优化
多模态搜索代理进行多跳推理仍面临根本性结构断层:现有流程独立构建训练数据、搜索环境和奖励信号,导致合成的结构化元数据被丢弃,环境依赖不可复现的外部引擎,强化学习奖励在轨迹层面依然稀疏。我们提出SearchEyes,以带类型的知识图谱为骨干,构建统一三者的模拟搜索世界。提出感知-知识链(PKC),在Wikidata5M的视觉-知识交集中采样受约束的多跳路径,保留每跳实体元数据,同时定义自洽的搜索世界与步骤级奖励锚点。进一步提出跳锚定策略优化(HaPO),复用这些锚点实现步骤级信用分配,无需独立训练的奖励模型。在六个多模态知识密集型基准上实验表明,SearchEyes在开源多模态搜索代理中达到最先进水平,SearchEyes-27B平均超越最强基线6.2分。
原文摘要 · Abstract (English)
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search environments, and reward signals independently, causing synthesized structural metadata to be discarded, environments to rely on irreproducible external engines, and RL rewards to remain sparse at the trajectory level. We present \textbf{SearchEyes}, which uses a typed knowledge graph as the backbone of a \emph{simulated search world} that unifies all three components. We propose \textbf{Perception-Knowledge Chains (PKC)} to sample constrained multi-hop paths over the visual-knowledge intersection of Wikidata5M, retaining hop-level entity metadata that simultaneously defines a self-contained search world and step-level reward anchors. We further propose \textbf{Hop-Anchored Policy Optimization (HaPO)}, which reuses these anchors for step-level credit assignment without a separately trained process reward model. Experiments on six multimodal knowledge-intensive benchmarks show that SearchEyes achieves state-of-the-art performance among open-source multimodal search agents, with SearchEyes-27B improving over the strongest open-source baseline by 6.2 points on average.%
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。