arXiv:2605.31354cs.AIcs.LG2026-05

弱模型协作时,共享记忆会放大幻觉,导致推理失败。

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

论文配图:Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents
图 1 · 摘自论文原文
  • 通过读-写-验证循环追踪视觉问答中的信息流
  • 4B-8B模型在多页文档中幻觉率上升37%
  • 适合研究低资源智能体协作可靠性的开发者

模块化视觉推理系统越来越多依赖共享工作内存进行多步协作,但低容量场景下中间状态演化的失败机制仍不明确。本文通过噪声累积视角研究弱模型(4B–8B参数量)协作推理的失败模式。提出CoSee审计框架,形式化读-写-验证循环以追踪文档视觉问答中的信息流。在跨多页文档、图表和网页的基准测试中,发现反直觉现象:简单共享空间常加剧幻觉而非缓解。识别出两种主导失败模式:噪声强化(无根据笔记被重复引用)与策略坍缩(上下文增加导致答案趋短且不明确)。基于成本-精度帕累托前沿分析表明,缺乏显式验证时,计算资源增加可能反而降低性能。结果表明,对资源受限代理而言,瓶颈不在推理深度,而在通信保真度,为可靠模块化设计提供了可追溯的诊断工具与机制基线。

原文摘要 · Abstract (English)

Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state evolution in low-capacity regimes remain underexplored. We study failure modes of collaborative reasoning with weak learners (4B--8B models) through the lens of noise accumulation. We introduce CoSee, an auditing framework that formalizes the read-write-verify loop to trace information flow in document visual question answering. Across multi-page, chart, and web-based benchmarks, we find a counter-intuitive degradation: naive shared workspaces often amplify hallucinations rather than resolve them. We identify two dominant failure modes: Noise Reinforcement, where ungrounded notes are reused as evidence, and Policy Collapse, where added context shifts the model toward under-specified, short-form answers. Using cost-accuracy Pareto frontiers, we show that increased compute can correlate negatively with performance without explicit verification. Our findings suggest that for resource-constrained agents, the bottleneck lies not in reasoning depth but in communication fidelity, providing trace-level diagnostics and a mechanistic baseline for reliable modular design.

视觉推理协作失效低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。