让研究型AI像人一样不断修正认知,避免错误累积。
VeriTrace: Evolving Mental Models for Deep Research Agents

- 构建认知图谱,通过三种反馈机制显式更新中间理解
- 在DeepResearch Bench上提升4.82个百分点,整体胜率增5.9%
- 适合需要高可靠推理的学术研究与复杂决策场景
深度研究代理面临海量、相互依赖且高度不确定的信息。现有系统虽探索了中间表示的演化形式,但将其演化完全交由大模型隐式推理完成。缺乏显式调控时,中间层易受低质量信息污染,错误沿依赖链传播,导致模型规模被用来弥补监管缺失。本文主张,代理的思维模型应通过持续反馈显式演化,以确保任务理解与现实对齐,并识别出三种调控回路:解释性更新、偏差反馈与模式重构。我们提出VeriTrace,一种实现这三重回路的认知图谱框架。采用匹配的Qwen3.5-27B作为基座模型,VeriTrace在DeepResearch Bench(DRB)洞察力指标上平均提升4.82个百分点(总体提升1.83个百分点),在DeepConsult上整体胜率提升5.9个百分点。使用Config-DeepSeek配置时,其在DRB上达成最强可复现开源结果。
原文摘要 · Abstract (English)
Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representations should look like, but leave their evolution to the LLM's implicit reasoning. Without explicit regulation, the intermediate layer is easily contaminated by mixed-quality information, and errors propagate along its dependencies, so model scale often ends up substituting for absent regulation. We argue that an agent's mental model should instead evolve through explicit feedback that continuously aligns task understanding with reality, and identify three regulatory loops: interpretive update, deviation feedback, and schema revision. We realise this in VeriTrace, a cognitive-graph framework that explicitly implements the three loops. Using matched Qwen3.5-27B backbones, VeriTrace improves over the strongest matched baseline by an average of 4.82 pp on DeepResearch Bench (DRB) Insight (1.83 pp Overall) and by 5.9 pp Overall win rate on DeepConsult. With Config-DeepSeek, it achieves the strongest reproducible open-source result on DRB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。