用结构化思维让智能体更懂多模态信息冲突
Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

- 构建动态多模态图结构,显式追踪知识演化过程
- 在多个数据集上平均提升17.2%准确率,超越现有最佳方法
- 适合需要处理复杂多源信息的深度研究场景
深度研究代理因其能大规模收集在线信息获取目标知识而受到越来越多关注,近期研究从纯文本信息检索转向多模态设置。然而,现有代理流程大多遵循证据累积模型,线性整合证据,缺乏对异构模态间矛盾信息的合理处理机制。为此,我们提出 Struct-Searcher,一种基于信念修正理论的结构化代理工作流,其在推理过程中显式维护一个不断演化的多模态结构图,实现具备冲突感知能力的多模态深度信息检索。在多个基准数据集和主干模型上的大量实验表明,Struct-Searcher具有(1)即插即用、模型无关特性,在五个不同主干模型上于 BrowseComp-VL 数据集上实现平均17.2%的相对准确率提升;(2)性能领先,持续优于当前最优的视觉语言模型(VLMs)和深度研究代理,在 MM-BrowseComp 上提升3.7%,在 HLE-VL 上提升1.5%,在 BrowseComp-VL 上提升0.7%(相较第二优方法)。
原文摘要 · Abstract (English)
Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with recent efforts shifting from purely text-based information seeking to multimodal settings. However, existing agentic workflows are largely aligned with evidence accumulation models, which linearly aggregate evidence and lack principled mechanisms for handling contradictory information across heterogeneous modalities. Towards this end, we propose Struct-Searcher, a structural agentic workflow grounded in belief revision theory that explicitly maintains an evolving multimodal structural graph throughout the reasoning process, enabling effective conflict-aware multimodal deep information seeking. Extensive experiments across multiple benchmark datasets and backbone models demonstrate that Struct-Searcher is (1) plug-and-play and model-agnostic, yielding an average relative accuracy improvement of 17.2% on BrowseComp-VL across five different backbones. (2) top-performing, consistently outperforming state-of-the-art vision-language models (VLMs) and deep research agents, with relative accuracy improvements of 3.7% on MM-BrowseComp, 1.5% on HLE-VL, and 0.7% on BrowseComp-VL over the second-best competing approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。