让智能体学会记忆和复用交互经验,提升多模态推理能力。
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
- 用事后推理抽象出原子化决策经验,构建可检索的经验库。
- 在视觉感知和复杂推理任务上超越现有基线,提升显著。
- 适合需要持续学习与策略优化的多模态智能体研究者。
研究智能体近年来在跨异构文本与视觉源的信息获取与整合方面取得显著进展。本文提出MuSEAgent,一种通过引入状态化经验来增强决策能力的多模态推理智能体。不同于依赖轨迹级检索的方法,我们提出一种状态化经验学习范式,通过事后推理将交互数据抽象为原子化决策经验,并组织成经过质量过滤的经验库,支持推理时的策略驱动经验检索。具体而言,MuSEAgent采用互补的广搜与深搜策略,实现跨多样化组合语义视角的动态多模态引导。大量实验表明,MuSEAgent在细粒度视觉感知与复杂多模态推理任务上均持续优于强基线方法。结果验证了状态化经验建模在提升多模态智能体推理能力方面的有效性。
原文摘要 · Abstract (English)
Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAgent, a multimodal reasoning agent that enhances decision-making by extending the capabilities of research agents to discover and leverage stateful experiences. Rather than relying on trajectory-level retrieval, we propose a stateful experience learning paradigm that abstracts interaction data into atomic decision experiences through hindsight reasoning. These experiences are organized into a quality-filtered experience bank that supports policy-driven experience retrieval at inference time. Specifically, MuSEAgent enables adaptive experience exploitation through complementary wide- and deep-search strategies, allowing the agent to dynamically retrieve multimodal guidance across diverse compositional semantic viewpoints. Extensive experiments demonstrate that MuSEAgent consistently outperforms strong trajectory-level experience retrieval baselines on both fine-grained visual perception and complex multimodal reasoning tasks. These results validate the effectiveness of stateful experience modeling in improving multimodal agent reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。