构建多模态长时序推理基准,测试模型从视觉语言音频中还原事件真相的能力。
MARPLE: A Benchmark for Long-Horizon Inference
- 设计模拟家庭场景下的多模态交互,通过逐步回放推断事件因果
- 人类表现优于传统方法和GPT-4,后者难以理解环境变化
- 验证视觉、语言、音频三模态均对推理有贡献,适合评估通用推理能力
重建过去事件需要跨长时间跨度的推理。为判断发生了什么,需结合世界知识与人类行为先验,并从视觉、语言和听觉等多种证据中推断。我们提出MARPLE,一个基于多模态证据评估长时序推理能力的基准。该基准包含在程序化生成环境中互动的智能体,支持视觉、语言和听觉刺激。受经典“谁犯案”故事启发,要求AI模型和人类参与者根据逐步回放还原实际发生过程,推断导致环境变化的元凶,并尽可能早地正确识别。实验表明,人类参与者在该任务上优于传统蒙特卡洛模拟方法和大语言模型基线(GPT-4)。相比人类,传统推理模型鲁棒性与性能较差,而GPT-4难以理解环境变化。我们分析了影响推理性能的因素并消融不同模态证据,发现三类模态对性能均有贡献。整体而言,本基准中的长时序多模态推理任务对当前模型仍具挑战性。
原文摘要 · Abstract (English)
Reconstructing past events requires reasoning across long time horizons. To figure out what happened, we need to use our prior knowledge about the world and human behavior and draw inferences from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchmark for evaluating long-horizon inference capabilities using multi-modal evidence. Our benchmark features agents interacting with simulated households, supporting vision, language, and auditory stimuli, as well as procedurally generated environments and agent behaviors. Inspired by classic ``whodunit'' stories, we ask AI models and human participants to infer which agent caused a change in the environment based on a step-by-step replay of what actually happened. The goal is to correctly identify the culprit as early as possible. Our findings show that human participants outperform both traditional Monte Carlo simulation methods and an LLM baseline (GPT-4) on this task. Compared to humans, traditional inference models are less robust and performant, while GPT-4 has difficulty comprehending environmental changes. We analyze what factors influence inference performance and ablate different modes of evidence, finding that all modes are valuable for performance. Overall, our experiments demonstrate that the long-horizon, multimodal inference tasks in our benchmark present a challenge to current models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。