用4D密室测试大模型的时间感知与跨模态主动感知能力
Evaluating Time Awareness and Cross-modal Active Perception of Large Models via 4D Escape Room Task
- 构建4D密室环境,融合声音触发、时间变化线索和位置依赖提示
- 模型在时间约束下跨模态整合能力弱,存在明显模态偏见
- 适合研究多模态推理、时空认知的AI开发者与研究人员
多模态大语言模型近年快速向集成视觉、语言和音频的通用模型发展。然而现有环境多聚焦2D或3D视觉上下文与视觉-语言任务,难以支持依赖时间的音频信号及选择性跨模态融合——不同模态可能互补或干扰,这对真实场景下的多模态推理至关重要。因此,模型能否在时变、不可逆条件下主动协调模态并推理仍待探索。为此,我们提出 extbf{EscapeCraft-4D},一个可定制的4D环境,用于评估通用模型的主动跨模态感知与时间意识。该环境包含基于触发的声音源、随时间消逝的证据以及位置依赖线索,要求智能体在时间压力下完成时空推理与主动多模态融合。基于此环境,我们构建基准评测多个先进模型的相关能力。评估结果表明,模型在模态偏见方面表现不佳,当前模型在时间约束下多模态融合能力存在显著差距。深入分析揭示了多种模态在复杂多模态推理环境中如何相互作用并共同影响决策。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently made rapid progress toward unified Omni models that integrate vision, language, and audio. However, existing environments largely focus on 2D or 3D visual context and vision-language tasks, offering limited support for temporally dependent auditory signals and selective cross-modal integration, where different modalities may provide complementary or interfering information, which are essential capabilities for realistic multimodal reasoning. As a result, whether models can actively coordinate modalities and reason under time-varying, irreversible conditions remains underexplored. To this end, we introduce \textbf{EscapeCraft-4D}, a customizable 4D environment for assessing selective cross-modal perception and time awareness in Omni models. It incorporates trigger-based auditory sources, temporally transient evidence, and location-dependent cues, requiring agents to perform spatio-temporal reasoning and proactive multimodal integration under time constraints. Building on this environment, we curate a benchmark to evaluate corresponding abilities across powerful models. Evaluation results suggest that models struggle with modality bias, and reveal significant gaps in current model's ability to integrate multiple modalities under time constraints. Further in-depth analysis uncovers how multiple modalities interact and jointly influence model decisions in complex multimodal reasoning environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。