arXiv:2511.15351cs.AIcs.CV2025-11被引 8

提出六能力协同的多模态智能体,可自主选择推理路径

Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration

  • 定义六类多模态推理能力并动态协调使用
  • 在Octopus-Bench上多数任务表现最优
  • 适合需要灵活决策的复杂多模态场景

现有多模态推理模型和框架存在根本性架构局限:大多缺乏人类般的自主探索能力——无论是直接推理、工具驱动的视觉探索、程序化视觉操作,还是内在的视觉想象。因此难以适应真实任务中不断变化的能力需求。而人类在处理此类任务时展现出互补的多种思维能力,现有方法却仅覆盖其中部分维度。受此启发,我们提出Octopus:一种六能力协同的多模态智能体推理新范式。我们定义了六项核心多模态推理能力,并据此构建了综合性评估基准Octopus-Bench。Octopus能在推理过程中自主探索,并根据当前状态动态选择最合适的能力建。实验表明,Octopus在Octopus-Bench的绝大多数任务上表现最佳,凸显了能力协调在多模态智能体推理中的关键作用。

原文摘要 · Abstract (English)

Existing multimodal reasoning models and frameworks suffer from fundamental architectural limitations: most lack the human-like ability to autonomously explore diverse reasoning pathways-whether in direct inference, tool-driven visual exploration, programmatic visual manipulation, or intrinsic visual imagination. Consequently, they struggle to adapt to dynamically changing capability requirements in real-world tasks. Meanwhile, humans exhibit a complementary set of thinking abilities when addressing such tasks, whereas existing methods typically cover only a subset of these dimensions. Inspired by this, we propose Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration, a new paradigm for multimodal agentic reasoning. We define six core capabilities essential for multimodal reasoning and organize a comprehensive evaluation benchmark, Octopus-Bench, accordingly. Octopus is capable of autonomously exploring during reasoning and dynamically selecting the most appropriate capability based on the current state. Experimental results show that Octopus achieves the best performance on the vast majority of tasks in Octopus-Bench, highlighting the crucial role of capability coordination in agentic multimodal reasoning.

多模态推理智能体能力协调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。