arXiv:2508.04361cs.AI2025-08被引 2

测试多模态模型在动态游戏中的跨模态推理能力,发现删掉感官信息反而能提升表现。

OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing

  • 设计五种游戏环境,强制模型融合视觉、听觉等多模态信息进行决策。
  • 六款主流模型在高保真记忆任务中表现超人,但策略规划能力严重不足。
  • 揭示感官冲突下模型易崩溃,且少用信息有时反而更有效。

尽管通用基础模型如Gemini和GPT-4o展现出出色的多模态能力,现有评估难以检验其在动态交互世界中的智能水平。静态基准缺乏主体性,而交互式基准则存在严重的模态瓶颈,通常忽略关键的听觉与时间线索。为弥合这一评估鸿沟,我们提出OmniPlay,一个诊断性基准平台,不仅用于评估,更旨在探测代理模型在全感官谱系下的融合与推理能力。基于模态互依的核心理念,OmniPlay包含五个游戏环境,系统构建了模态协同与冲突场景,迫使代理进行真实的跨模态推理。对六款领先多模态模型的全面评估显示显著二分:它们在高保真记忆任务上表现超人,但在需稳健推理与战略规划的任务中存在系统性失败。我们证明这种脆弱性源于不稳定的融合机制,在模态冲突下导致性能灾难性下降,并揭示一个反直觉的‘少即是多’悖论——移除感官信息反而可能提升表现。研究提示,通往鲁棒通用人工智能需超越规模扩展,聚焦于协同融合机制的专门研究。平台已开放匿名评审:https://github.com/fuqingbie/omni-game-benchmark。

原文摘要 · Abstract (English)

While generalist foundation models like Gemini and GPT-4o demonstrate impressive multi-modal competence, existing evaluations fail to test their intelligence in dynamic, interactive worlds. Static benchmarks lack agency, while interactive benchmarks suffer from a severe modal bottleneck, typically ignoring crucial auditory and temporal cues. To bridge this evaluation chasm, we introduce OmniPlay, a diagnostic benchmark designed not just to evaluate, but to probe the fusion and reasoning capabilities of agentic models across the full sensory spectrum. Built on a core philosophy of modality interdependence, OmniPlay comprises a suite of five game environments that systematically create scenarios of both synergy and conflict, forcing agents to perform genuine cross-modal reasoning. Our comprehensive evaluation of six leading omni-modal models reveals a critical dichotomy: they exhibit superhuman performance on high-fidelity memory tasks but suffer from systemic failures in challenges requiring robust reasoning and strategic planning. We demonstrate that this fragility stems from brittle fusion mechanisms, which lead to catastrophic performance degradation under modality conflict and uncover a counter-intuitive "less is more" paradox, where removing sensory information can paradoxically improve performance. Our findings suggest that the path toward robust AGI requires a research focus beyond scaling to explicitly address synergistic fusion. Our platform is available for anonymous review at https://github.com/fuqingbie/omni-game-benchmark.

多模态游戏评测推理能力AGI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。