提出虚拟场景推理新基准,测试模型在无实时数据下的三维理解能力
Hypo3D: Exploring Hypothetical Reasoning in 3D
- 构建基于假设变化的3D视觉问答框架,模拟非实时场景推理
- 包含700个室内场景、14885组问答对,覆盖7727次场景变化
- 发现当前大模型在方向性推理和运动变化任务中表现远低于人类
视觉语言基础模型的发展推动了机器在三维场景推理中接近人类能力。现有三维推理基准假设可实时访问场景,但频繁更新成本过高。为此,我们提出假设性三维推理(Hypo3D),一个评估模型在无法获取实时场景数据时推理能力的基准。模型需根据给定的变化描述想象场景状态后再进行推理。Hypo3D被设计为3D视觉问答(VQA)基准,涵盖700个室内场景中的7727次上下文变化,生成14885组问答对。所有场景建立基于锚点的世界坐标系,确保方向性术语在上下文与问题中具有一致全局参考。大量实验表明,最先进基础模型在假设性场景推理中表现不佳,与人类相比存在显著差距,尤其在涉及移动变化和方向性推理的任务中。即使上下文变化与问题无关,模型仍常错误调整答案。
原文摘要 · Abstract (English)
The rise of vision-language foundation models marks an advancement in bridging the gap between human and machine capabilities in 3D scene reasoning. Existing 3D reasoning benchmarks assume real-time scene accessibility, which is impractical due to the high cost of frequent scene updates. To this end, we introduce Hypothetical 3D Reasoning, namely Hypo3D, a benchmark designed to evaluate models' ability to reason without access to real-time scene data. Models need to imagine the scene state based on a provided change description before reasoning. Hypo3D is formulated as a 3D Visual Question Answering (VQA) benchmark, comprising 7,727 context changes across 700 indoor scenes, resulting in 14,885 question-answer pairs. An anchor-based world frame is established for all scenes, ensuring consistent reference to a global frame for directional terms in context changes and QAs. Extensive experiments show that state-of-the-art foundation models struggle to reason in hypothetically changed scenes. This reveals a substantial performance gap compared to humans, particularly in scenarios involving movement changes and directional reasoning. Even when the context change is irrelevant to the question, models often incorrectly adjust their answers. Project website: https://matchlab-imperial.github.io/Hypo3D/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。