arXiv:2602.22897cs.AIcs.CL2026-02被引 21

构建跨模态智能体评测基准,推动通用人工智能助手发展

OmniGAIA: Towards Native Omni-Modal AI Agents

  • 基于多模态事件图生成复杂任务,支持视频/音频/图像跨模态推理
  • 提出OmniAtlas模型,通过回溯引导探索提升工具使用能力
  • 适合研究多模态推理与智能体系统的设计者和开发者

人类智能天然融合视觉、听觉与语言的全模态感知,并结合复杂推理与工具使用来与世界互动。然而,现有多模态大模型主要局限于双模态交互(如视觉-语言),缺乏通用智能助手所需的统一认知能力。为此,我们提出OmniGAIA,一个全面的基准评测体系,用于评估在视频、音频和图像模态下需深度推理与多轮工具执行的任务表现。该基准通过新颖的全模态事件图方法构建,生成源自真实数据的复杂多跳查询,要求跨模态推理与外部工具集成。同时,我们提出OmniAtlas,一种在工具集成推理范式下具备主动全模态感知能力的原生全模态基础智能体。其训练采用回溯引导树探索策略合成轨迹,并通过OmniDPO进行细粒度错误修正,显著增强现有开源模型的工具使用能力。本工作为下一代面向现实场景的原生全模态智能助手迈出了关键一步。

原文摘要 · Abstract (English)

Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However, current multi-modal LLMs are primarily confined to bi-modal interactions (e.g., vision-language), lacking the unified cognitive capabilities required for general AI assistants. To bridge this gap, we introduce OmniGAIA, a comprehensive benchmark designed to evaluate omni-modal agents on tasks necessitating deep reasoning and multi-turn tool execution across video, audio, and image modalities. Constructed via a novel omni-modal event graph approach, OmniGAIA synthesizes complex, multi-hop queries derived from real-world data that require cross-modal reasoning and external tool integration. Furthermore, we propose OmniAtlas, a native omni-modal foundation agent under tool-integrated reasoning paradigm with active omni-modal perception. Trained on trajectories synthesized via a hindsight-guided tree exploration strategy and OmniDPO for fine-grained error correction, OmniAtlas effectively enhances the tool-use capabilities of existing open-source models. This work marks a step towards next-generation native omni-modal AI assistants for real-world scenarios.

多模态智能体推理工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。